Shuichiro Ogawa
日本語

Notes · updated 2026-09-27

Has RAG Fallen Out of Use? What Replaced It, and What Remains

The rumor that RAG is no longer used does not hold as a description of actual use. Enterprise RAG adoption rose from 31% in 2023 to 51% in 2024 (Menlo, 600 respondents), and the 2025 survey still ranks it second after prompt design.

Contents (11)
  1. The Rumor Came From Codebases
  2. With Code, the Index Loses
  3. The Obituary Writer Still Runs Retrieval
  4. T1v Vendor Primary: RAG Is Being Rebuilt, Not Retired
  5. T2 Public Bodies and Research: Adoption Figures, and What the Standards Watch
  6. How Far Does Long Context Go?
  7. Context Engineering Contains RAG as a Part
  8. Practices Available as of 2026
  9. Jev Does Not Retrieve, but It Enters the Judgment Stages of Retrieval
  10. Recent Major Updates (Chronological)
  11. Reading the Tiers

The Rumor Came From Codebases

The phrase “RAG is dead” was repeated many times in 2025. Trace it back and nearly every instance leads to coding agents.

On 2025-05-27, Cline’s official blog published “Why Cline Doesn’t Index Your Codebase.” The same week, LlamaIndex published a post titled “RAG is dead, long live agentic retrieval.” On 2025-09-11, Jason Liu wrote “Why I Stopped Using RAG for Coding Agents (And You Should Too),” based on a talk by Nik Pash, Head of AI at Cline. On 2025-10-01, Fintool founder Nicolas Bustamante published “The RAG Obituary: Killed by Agents, Buried by Context Windows.” On 2026-05-14, Anthropic’s official blog explained why Claude Code searches with grep rather than an embedding index.

Each of these is about one particular kind of source: a codebase. Whether the rumor is true depends on whether what happened with code carries over to other sources.

With Code, the Index Loses

Anthropic’s explanation centers on staleness. “embedding pipelines can’t keep up with active engineering teams. By the time a developer queries the index, it reflects the codebase as it previously existed weeks, days, or even hours before.” In a codebase that thousands of people keep committing to, the code changes before the embeddings can be rebuilt. Claude Code instead “traverses the file system, reads files, uses grep to find exactly what it needs.” As a result, “There’s no embedding pipeline or centralized index to maintain.”

Cline gave three reasons. Chunking cuts the logic of code apart (“you’re literally tearing apart its logic”). An index is by definition a copy of one moment, and the code drifts away from it (“An index, by definition, is a snapshot frozen in time”). And placing embeddings outside means sending out the code that is the source of a company’s advantage.

In an interview with Claude Code’s creator Boris Cherny, the interviewer Gergely Orosz summarizes the history. The team tried several approaches, including local vector databases, and each had downsides: stale indexes and permission complexity. “Plain glob and grep, driven by the model, beat everything.” This is the interviewer’s summary, not Cherny’s own words.

The reason grep works for code lies in the nature of the source. Function and type names can be searched verbatim, and following references leads to the related places. There is little need to search by closeness of meaning, and the content changes by the hour. Sources such as internal policies, papers, or case law, where wording varies and content changes slowly, do not meet these conditions as they stand.

The Obituary Writer Still Runs Retrieval

Read by its title alone, Bustamante’s post declares the end of RAG. Its conclusion is “In hindsight, RAG will look like training wheels. Useful, necessary, but temporary.” That is a prediction. The same post says Fintool currently runs hybrid search combining chunking, embeddings, and BM25. A full move to agentic search appears only as the future the author expects.

Voices came from the other side too. Hamel Husain wrote “I’m tired of hearing ‘RAG is dead.’” and organized an open seven-part series with Ben Clavié (2025-07-12). Chroma’s Jeff Huber set the word aside before the question of life or death. “We never use the term rag. I hate the term rag.” The word Huber uses instead is context engineering, defined as “the job of figuring out what should be in the context window for any given LLM generation step.”

The LlamaIndex post, too, carries the in-text heading “Naive RAG is dead, agentic retrieval is the future.” What is declared dead is the naive setup, not retrieval itself. The text continues: “now agentic strategies are table stakes.”

T1v Vendor Primary: RAG Is Being Rebuilt, Not Retired

The major providers added RAG products throughout 2025.

On 2025-11-06, Google built File Search into the Gemini API. The official description is “a fully managed RAG system built directly into the Gemini API.” It handles chunking, embedding, and injection into context automatically, and returns citations. Only the embeddings for initial indexing are billed, at $0.15 per million tokens; storage and query-time embeddings are free.

OpenAI’s File Search is a hosted tool in the Responses API and combines “semantic and keyword search.” Deep research (2025-02-02) arrived as an agent that carries out multi-step research on the internet.

Microsoft added agentic retrieval to Azure AI Search. It is defined as “a multi-query pipeline designed for complex questions.” It decomposes a compound question into subqueries, runs them in parallel, and semantically reranks each. The billing unit changed from queries to tokens. It is not uniformly generally available: the extractive minimal configuration went GA in the 2026-04-01 REST API, while LLM-based query planning and answer synthesis remain in preview.

Graph-based retrieval also entered products. AWS made GraphRAG in Bedrock Knowledge Bases generally available on 2025-03-07, and Microsoft Research added a 2025-06 note that LazyGraphRAG had been integrated into its research platform and Azure Local.

Anthropic ships the surroundings of retrieval as components. The Citations API ties claims in a response to character ranges or pages of the documents, and the web search tool performs cited searches at $10 per 1,000 searches. Google’s check grounding is an API that judges how far an answer is supported by a given set of facts.

What this line-up shows is not the retirement of RAG but vendors taking over the assembly of retrieval. Chunking and embedding disappear inside managed products, and the units of design that remain visible are how many times to search and what to return as citations.

T2 Public Bodies and Research: Adoption Figures, and What the Standards Watch

The only source found that shows the adoption trend with a stated methodology is Menlo Ventures’ survey. The 2024 edition (600 US IT decision-makers, 2024-09-24 to 10-08) writes “RAG … now dominates at 51% adoption, a dramatic rise from 31% last year.” The 2025 edition (495 respondents, 2025-11-07 to 25) gives no percentage, but ranks prompt design first among customization techniques and RAG second. The same 2025 edition says only 16% of enterprise and 27% of startup deployments qualify as true agents. Menlo is a VC with portfolio companies, so the figures are read as the publisher’s claims.

The other surveys say nothing about RAG. LangChain’s survey (1,340 responses, 2025-11 to 12) puts agents in production at 57.3% (51% the year before) but has no figure for RAG alone. The body of a16z’s CIO survey (2025-06) mentions neither RAG, retrieval, nor context engineering. Where Gartner’s Hype Cycle placed RAG could not be read, since the original returned 403 and only third-party summaries were available, so this note does not use it.

The standards side has begun to treat RAG as an attack surface. OWASP added “LLM08:2025 Vector and Embedding Weaknesses” to the 2025 Top 10 for LLM. It names the “risk of context leakage between users or queries” where many users share a vector database, and lists “permission-aware vector and embedding stores” as a mitigation. NIST’s NCCoE published on 2025-07-31 its record of building a RAG chatbot in-house (IR 8579, initial public draft), covering prompt injection, hallucination, data exposure, and unauthorized access. AIST’s Generative AI Quality Management Guideline (2025-05-26) names RAG components explicitly as objects of quality management.

All of these documents assume that RAG is widely deployed and set about closing its weaknesses. None was found that treats it as an abandoned setup.

How Far Does Long Context Go?

One view holds that once context windows reach a million tokens, retrieval becomes unnecessary. Anthropic itself grants this, with a condition. “If your knowledge base is smaller than 200,000 tokens (about 500 pages of material), you can just include the entire knowledge base in the prompt.” At that size, the 2024-09 Contextual Retrieval post recommended putting everything in and pairing it with prompt caching.

A year later, the same Anthropic took the opposite fact as its premise. “as the number of tokens in the context window increases, the model’s ability to accurately recall information from that context decreases.” Whether something fits and whether it can be read are separate questions.

Chroma’s Context Rot study (2025-07-14) measured this across 18 models. Its conclusion: “model performance degrades as input length increases, often in surprising and non-uniform ways.” The less the target information resembled the wording of the question, the more performance fell as the input grew. This suggests that questions involving paraphrase are not solved by putting everything into a long context.

Drew Breunig sorted the failures of long context into four kinds (2025-06-22). Poisoning, where an error enters the context and is referred to repeatedly; distraction, where the model leans on the context and stops using what it learned in training; confusion, where superfluous content leads to a low-quality response; and clash, where new information conflicts with what is already there. Breunig cites a study (arXiv:2505.06120) in which splitting a prompt across several turns produced an average drop of 39%.

Pinecone writes, on cost grounds, that long context makes each query more expensive and slower. Pinecone sells a vector database, however, and offers no quantitative comparison.

Long context can stand in for retrieval when the knowledge is small and stable. With large knowledge, or with questions whose wording varies, the work of choosing what to put in remains.

Context Engineering Contains RAG as a Part

In June 2025, context engineering spread as the successor term to prompt engineering. Tobi Lütke wrote “the art of providing all the context for the task to be plausibly solvable by the LLM,” and Karpathy wrote “the delicate art and science of filling the context window with just the right information for the next step.” The genealogy of the term is set out in The Genealogy of Generative AI Engineering and Its Assessment from a Design Perspective.

Inside this term, RAG did not disappear; it became one move among several. Lance Martin (LangChain) divides context operations into write, select, compress, and isolate, and retrieval sits as one form of select. Anthropic states the goal as “find the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome,” and names just-in-time loading as a way to reach it. The agent holds lightweight identifiers such as file paths and stored queries, and loads the content with tools when it becomes needed.

The same post names Claude Code as an example of the hybrid. “CLAUDE.md files are naively dropped into context up front, while primitives like glob and grep allow it to navigate its environment and retrieve files just-in-time.” The small amount that is always needed goes in first; the rest is searched for at run time. Searching with grep is also retrieval, in the sense of choosing what enters the context. What changed is the move from a setup that searches once in advance to one where the agent calls search repeatedly while making judgments.

Practices Available as of 2026

Only practices that primary sources give with conditions are listed here.

  • If the knowledge is under about 200,000 tokens and stable, do not retrieve: put all of it in and use prompt caching (Anthropic). Recall still falls as context grows, so avoid filling it to the limit.
  • For large document sets, add context to chunks, combine with BM25, and re-rank: in Anthropic’s own evaluation, the top-20 retrieval failure rate fell from 5.7% to 3.7% with contextual embeddings, 2.9% with BM25 added, and 1.9% with re-ranking (metric 1 − recall@20, sample size not disclosed). Take 150 candidates in the first pass, re-rank down to 20, and pass 20 rather than 5 or 10 to the model.
  • For fast-changing sources addressable by identifiers, let the agent search: codebases are the typical case. Exploring at run time is slower than retrieving in advance, so Anthropic recommends a hybrid that loads what is always needed up front.
  • Decompose compound questions before retrieving: Azure’s agentic retrieval turned this into a product. Billing becomes per token, so the cost per question moves with its complexity.
  • Return answers tied to their sources: as with the Citations API, search results blocks, and check grounding, include the correspondence between claims and evidence in the output so that it can be checked afterwards.
  • Do not trust retrieval results: partition vector databases by user permissions, and check knowledge sources for injection and poisoning (OWASP LLM08).

Jev Does Not Retrieve, but It Enters the Judgment Stages of Retrieval

TypeSafe AI’s Jev (TypeSafe AI's Jev: What It Guarantees and What It Does Not) generates no text and returns only typed judgments and probabilities. It retrieves nothing on its own. Even so, the official cookbooks contain four examples that use it as a component of RAG.

The first is re-ranking. On 40 queries from the case-law corpus CLERC, BM25 narrowed 3,565 passages to 30 candidates, which were then reordered with the Noul question “Could this candidate passage be from that cited precedent?” The share of queries with the correct passage in first place rose from 5% to 18%, and within the top ten from 38% to 62%, for $0.0645 across 1,200 calls (jev-1.12). The only comparison is against BM25’s own order; there is no comparison with embedding retrieval or dedicated re-ranking models. The 30 candidates contained the correct passage for all 40 queries, and the vendor itself states the limit: “Re-ranking only ever sees the passages that make the shortlist.”

The second is passage screening, which places four questions between retrieval and generation. Is the passage relevant, does it contain usable evidence, does it contradict the premise of the question, and does it try to manipulate the answering system? Thresholds route each passage to exclusion, evidence, or contradiction, and contradicting evidence is passed to the generator as a separate block. The sample is 81 passages of Supabase authentication docs and six queries, and on injection detection the vendor states plainly: “Nothing here is a security boundary.”

The third is citation checking. It looks for the quoted text by string matching; if absent the citation is fabricated, and if present a Choice question judges whether the section supports, contradicts, or is unrelated to the claim, with confidence below 0.8 sent to human review. The fourth is line-level semantic search: each line of a document gets an ID, a Choice question ranks them, and a Noul question judges at the same time whether the answer is in the document at all. Because Choice probabilities sum to one, some line ranks first even when no answer exists. The Noul question closes that gap.

What the four share is that Jev is given the judgments that sit outside retrieval. BM25 or embeddings gather the candidates, and Jev judges whether a result is relevant, whether it counts as evidence, and whether it can be trusted. When Judgment Becomes a Function Call, What Moves in Design argued that applying criteria descends into components, while the work of writing the criteria remains. The judgment stages of RAG are a typical case of that application, and the text of the relevance question plays the role of the criterion. The rise from 5% to 18% is the result of writing a single question, and the vendor notes that a real application would ask several.

The dependence also runs the other way. Jev itself loses accuracy in long contexts. Its state holds around 32,000 tokens according to the official docs (another page says 64k, and the docs disagree with themselves); an independent implementation reports “Jev suffers from context rot,” and the vendor instructs “Include only the context relevant to the current questions” (How to Use Jev, and Where Not To). Using Jev requires choosing, one step earlier, what to hand it. In other words, Jev does not make retrieval unnecessary; it presupposes that retrieval exists.

The typia maintainer chose Jev over embedding search to select, turn by turn, the functions needed out of several hundred (How to Use Jev, and Where Not To). It lines up an independent Noul for each candidate, and when there are too many, selects a group first and then individual functions. It is an example where, if the candidate set is small and fixed, one can skip retrieval and select by judgment alone.

Recent Major Updates (Chronological)

  • 2024-09-19 Anthropic Contextual Retrieval. Recommends including everything below 200,000 tokens
  • 2024-11-18 OWASP Top 10 for LLM, 2025 edition. Adds LLM08 Vector and Embedding Weaknesses
  • 2024-11-20 Menlo survey. RAG adoption 51% (31% the year before)
  • 2025-01 Anthropic Citations API
  • 2025-02-02 OpenAI Deep research
  • 2025-03-07 AWS Bedrock Knowledge Bases GraphRAG GA
  • 2025-05-07 Anthropic web search tool (web fetch added 2025-09-10)
  • 2025-05-26 AIST Generative AI Quality Management Guideline
  • 2025-05-27 Cline “Why Cline Doesn’t Index Your Codebase”
  • 2025-05-29 LlamaIndex “RAG is dead, long live agentic retrieval”
  • 2025-06-19 to 27 Lütke, Karpathy, Martin, Willison, and Breunig discuss context engineering
  • 2025-07-12 Hamel Husain and Ben Clavié, “Stop Saying RAG Is Dead”
  • 2025-07-14 Chroma Context Rot (18 models)
  • 2025-07-31 NIST NCCoE IR 8579 initial public draft
  • 2025-09-11 Jason Liu, “Why I Stopped Using RAG for Coding Agents”
  • 2025-09-29 Anthropic “Effective context engineering for AI agents”
  • 2025-10-01 Bustamante, “The RAG Obituary”
  • 2025-11-06 Google Gemini API File Search
  • 2025-12-09 Menlo 2025 edition. Prompt design first among customization techniques, RAG second
  • 2026-03-04 The Pragmatic Engineer interview with Boris Cherny
  • 2026-04-01 Minimal configuration of Azure AI Search agentic retrieval goes GA
  • 2026-05-14 Anthropic “How Claude Code works in large codebases”

Reading the Tiers

  • T1v vendor primary: official material from Anthropic, Google, OpenAI, Microsoft, AWS, LlamaIndex, Pinecone, Cohere, and TypeSafe, treated as the primary authority on features, prices, and design rationale. The Contextual Retrieval figures are an in-house evaluation with undisclosed sample size. LlamaIndex’s title claim, Cohere’s “state-of-the-art,” and Pinecone’s conclusion that RAG is not obsolete were separated as marketing in the corpus.
  • T2 public bodies and research: OWASP, NIST, and AIST are primary sources that treat RAG as an attack surface and an object of quality management. Adoption figures come from a single series, Menlo’s, and the publisher being a VC with portfolio companies needs to be discounted. Gartner and McKinsey could not be reached, and a search summary claiming “McKinsey reports RAG at 51%” was judged a misattribution of Menlo’s figure and not used.
  • T3 individual opinion: Karpathy, Willison, Breunig, Martin, Huber, Husain, Bustamante, and Liu were confirmed as authorities, but all of this is opinion. Huber and Bustamante are also parties with stakes in retrieval infrastructure or retrieval products. Cherny’s post on X could not be reached, so only the interviewer’s summary was used.

The sources agree on the direction: naive one-shot retrieval has receded, and retrieval has become a tool for agents. Where they split is whether to call that “the death of RAG” or “the evolution of RAG,” which is mainly a difference in how the word is defined.

Related notes: TypeSafe AI's Jev: What It Guarantees and What It Does Not, How to Use Jev, and Where Not To, When Judgment Becomes a Function Call, What Moves in Design, The Genealogy of Generative AI Engineering and Its Assessment from a Design Perspective, Agentic Coding: The Current State of Orchestration Patterns (2026), Recent Currents in LLM-as-a-Judge: A Literature Map of 53 Core Studies and a Standalone Chapter on Creativity Evaluation.

Unverified Items

  • [要一次検証: 取得不能 x.com 402, xcancel 451, not on threadreaderapp] The verbatim text and date of Boris Cherny’s post on X, “Early versions of Claude Code used RAG + a local vector db.” The body uses only the interviewer’s summary.
  • [要一次検証: 取得不能 gartner.com press release and article page returned 403; the original document is subscriber-only] The positions of RAG, AI agents, and vector databases on Gartner’s Hype Cycle for AI 2025.
  • [要一次検証: 取得不能 mckinsey.com page and PDF timed out twice] What McKinsey’s State of AI 2025 says about RAG.
  • [要一次検証: 本文未到達 pp.21–47 not read] Whether the latter part of §3 of NIST AI 600-1 mentions retrieval or RAG. Pages 1–20 do not.
  • [要一次検証: 本文未到達 WebFetch 403, search snippet only] The body of OpenAI’s Deep research announcement.
  • [要一次検証: 本文未到達 fetched content was mostly a summary] The verbatim “lost in the middle” wording in Pinecone’s post.
  • [要一次検証: 未試行 checking the original paper was out of budget] The original study behind Breunig’s “average drop of 39%” (arXiv:2505.06120). The body attributes it to Breunig’s citation.
  • [要確認: 記載なし no date on each page] Publication dates of the four TypeSafe cookbooks, OpenAI File Search, and Anthropic’s search results block.
  • [要確認: 未試行 time budget] How UK AISI, the main volume of the Stanford HAI AI Index, and Japan AISI’s evaluation guide treat RAG.

References

All accessed 2026-09-27.

Vendor primary (T1v)

Public bodies and research (T2)

Individual opinion (T3)


Author: Shuichiro Ogawa (Design Researcher / Consultant) About me →