Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

RAG Checklist Targets Retrieval Failures

RAG Checklist Targets Retrieval Failures

DEV.to·Wednesday, August 26, 2026
  • •James Anderson says roughly 73% of RAG failures occur in retrieval, not answer generation
  • •Checklist separates offline indexing from online query handling under an under ~3 seconds latency target
  • •Recommended RAG recipe retrieves ~20 chunks, reranks to ~5, then sends 3–5 to the LLM
  • •James Anderson says roughly 73% of RAG failures occur in retrieval, not answer generation
  • •Checklist separates offline indexing from online query handling under an under ~3 seconds latency target
  • •Recommended RAG recipe retrieves ~20 chunks, reranks to ~5, then sends 3–5 to the LLM
  • •James Anderson says roughly 73% of RAG failures occur in retrieval, not answer generation
  • •Checklist separates offline indexing from online query handling under an under ~3 seconds latency target
  • •Recommended RAG recipe retrieves ~20 chunks, reranks to ~5, then sends 3–5 to the LLM
  • •James Anderson says roughly 73% of RAG failures occur in retrieval, not answer generation
  • •Checklist separates offline indexing from online query handling under an under ~3 seconds latency target
  • •Recommended RAG recipe retrieves ~20 chunks, reranks to ~5, then sends 3–5 to the LLM

James Anderson published a RAG deployment checklist on August 25, 2026, arguing that many wrong answers blamed on LLMs begin earlier in the retrieval pipeline. He says industry analysis in 2026 puts retrieval at roughly 73% of RAG failures, because the model often summarizes the context it receives even when that context contains the wrong chunks. Anderson frames production RAG as two separate paths: an offline indexing path that parses, cleans, chunks, enriches, embeds and writes documents to vector and keyword indexes, and an online query path that rewrites queries, retrieves candidates, reranks, assembles citations, generates answers and logs traces under a latency target of under ~3 seconds end to end.

Anderson identifies chunking as a common silent failure point, warning that a default split of every 1,000 characters with 100 overlap can cut sentences, tables and code in ways that produce relevant-looking but incomplete evidence. He recommends structure-aware splitting around document headings, code functions or table rows, and semantic chunking (splitting when meaning shifts) when extra compute is justified. Each chunk, he says, should be able to answer a question on its own; chunks that only make sense beside their neighbors are too fragmented, while oversized chunks dilute the signal.

The checklist says teams should embed chunks with surrounding context rather than body text alone, such as a heading, short document summary or one-line description, and keep metadata like author, date, source, section, product version, document type and access level attached. Anderson also says vector search alone is a major mistake: semantic retrieval handles vague phrases such as login problems, but exact strings such as `ERR_SSL_PROTOCOL_ERROR`, `WX-4200` or a function name need keyword search. He recommends hybrid search using BM25 or full-text search plus vector search, fused with Reciprocal Rank Fusion (RRF), citing 2024–2026 benchmarks including BEIR and MTEB where BM25 plus dense embeddings fused with RRF beats either approach alone.

Anderson says raw user questions often need query transformation before retrieval because conversational or compound wording may not match precise indexed statements. The patterns he lists are query rewriting or expansion, HyDE (generating a hypothetical answer for retrieval), step-back prompting and decomposition into separate subqueries. For ordering results, he recommends retrieving a broad set with hybrid search, then using a cross-encoder reranker that scores each query and chunk together. He says reranking commonly adds 5–15 points of MRR and has pushed nDCG@10 from ~0.13 to ~0.40 on some reasoning-heavy benchmarks, roughly 3x from reordering the same candidates. His practical recipe is to retrieve ~20 via hybrid search, rerank to ~5, and send 3–5 chunks to the LLM.

The final sections warn that prompt assembly can still damage retrieval quality if teams send too many chunks, bury the strongest evidence in the middle, omit citations or treat million-token context windows as a substitute for retrieval. Anderson describes agentic RAG as useful for nuanced multi-hop questions because retrieval moves into a ReAct-style loop where the model can retrieve, inspect results, rewrite the query and retrieve again. He says agentic RAG costs more in calls, latency and nondeterminism, and is best reserved for genuinely complex or high-stakes retrieval such as legal, medical, financial or multi-hop work after single-pass retrieval is solid. He closes by urging teams to measure retrieval and generation separately with recall, rank, nDCG, MRR and faithfulness checks, because end-to-end impressions can hide whether a failure came from retrieval or answer generation.

James Anderson published a RAG deployment checklist on August 25, 2026, arguing that many wrong answers blamed on LLMs begin earlier in the retrieval pipeline. He says industry analysis in 2026 puts retrieval at roughly 73% of RAG failures, because the model often summarizes the context it receives even when that context contains the wrong chunks. Anderson frames production RAG as two separate paths: an offline indexing path that parses, cleans, chunks, enriches, embeds and writes documents to vector and keyword indexes, and an online query path that rewrites queries, retrieves candidates, reranks, assembles citations, generates answers and logs traces under a latency target of under ~3 seconds end to end.

Anderson identifies chunking as a common silent failure point, warning that a default split of every 1,000 characters with 100 overlap can cut sentences, tables and code in ways that produce relevant-looking but incomplete evidence. He recommends structure-aware splitting around document headings, code functions or table rows, and semantic chunking (splitting when meaning shifts) when extra compute is justified. Each chunk, he says, should be able to answer a question on its own; chunks that only make sense beside their neighbors are too fragmented, while oversized chunks dilute the signal.

The checklist says teams should embed chunks with surrounding context rather than body text alone, such as a heading, short document summary or one-line description, and keep metadata like author, date, source, section, product version, document type and access level attached. Anderson also says vector search alone is a major mistake: semantic retrieval handles vague phrases such as login problems, but exact strings such as `ERR_SSL_PROTOCOL_ERROR`, `WX-4200` or a function name need keyword search. He recommends hybrid search using BM25 or full-text search plus vector search, fused with Reciprocal Rank Fusion (RRF), citing 2024–2026 benchmarks including BEIR and MTEB where BM25 plus dense embeddings fused with RRF beats either approach alone.

Anderson says raw user questions often need query transformation before retrieval because conversational or compound wording may not match precise indexed statements. The patterns he lists are query rewriting or expansion, HyDE (generating a hypothetical answer for retrieval), step-back prompting and decomposition into separate subqueries. For ordering results, he recommends retrieving a broad set with hybrid search, then using a cross-encoder reranker that scores each query and chunk together. He says reranking commonly adds 5–15 points of MRR and has pushed nDCG@10 from ~0.13 to ~0.40 on some reasoning-heavy benchmarks, roughly 3x from reordering the same candidates. His practical recipe is to retrieve ~20 via hybrid search, rerank to ~5, and send 3–5 chunks to the LLM.

The final sections warn that prompt assembly can still damage retrieval quality if teams send too many chunks, bury the strongest evidence in the middle, omit citations or treat million-token context windows as a substitute for retrieval. Anderson describes agentic RAG as useful for nuanced multi-hop questions because retrieval moves into a ReAct-style loop where the model can retrieve, inspect results, rewrite the query and retrieve again. He says agentic RAG costs more in calls, latency and nondeterminism, and is best reserved for genuinely complex or high-stakes retrieval such as legal, medical, financial or multi-hop work after single-pass retrieval is solid. He closes by urging teams to measure retrieval and generation separately with recall, rank, nDCG, MRR and faithfulness checks, because end-to-end impressions can hide whether a failure came from retrieval or answer generation.

Read original (English)·Aug 25, 2026
#rag#retrieval#hybrid search#bm25#reciprocal rank fusion#reranking#hyde#graphrag#mrr#ndcg