Your coding agent finds code by grepping and reading files, not by embeddings. Anthropic says that works. Cursor says semantic search added accuracy. Here is how to decide which your own app needs.
You are using Cursor, Claude Code or Lovable, and you read that real AI apps need a vector database. Then you notice your coding agent seems to find code just fine without one. Which is it?
Both are true, for different jobs. This page explains the two approaches in plain English, shows what the two biggest vendors say, and ends with rules you can apply to your own app. New to the vocabulary? Start with what a context window is and context engineering for vibe coders.
What is agentic search?
Agentic search means the AI agent searches for itself, with tools, while it works. It runs grep to find a word, lists files by name, opens one, reads it, then decides what to look at next. Nothing is indexed ahead of time.
Anthropic calls the pattern "just in time" context. Its September 2025 engineering post says such agents "maintain lightweight identifiers (file paths, stored queries, web links, etc.) and use these references to dynamically load data into context at runtime using tools."
Grep is a command that finds lines matching a pattern. In Claude Code, the docs list a Search tool category that can "find files by pattern, search content with regex, explore codebases." On macOS, Linux and WSL, the tools reference says Claude searches with find and grep through the Bash tool.
What is RAG, and where does the vector database come in?
RAG stands for retrieval-augmented generation. You prepare your content ahead of time, then fetch the relevant pieces and paste them into the prompt before the AI answers. A vector database is where the prepared pieces live.
Anthropic's Contextual Retrieval post lists the steps. Break the content into chunks of "usually no more than a few hundred tokens." Use an embedding model to turn each chunk into a vector, a list of numbers that encodes meaning. Store them in a vector database. At question time, find the chunks closest in meaning to the question and add them to the prompt.
You can try this in a database you may already use. Supabase's docs say vectors there are enabled via pgvector, "a Postgres extension for storing and querying vectors." One more fact: the Claude API docs say "Anthropic does not offer its own embedding model," and point to Voyage AI as one provider. For a hands-on version, see how to chat with your documents locally.
What does Anthropic say about when to use each?
Anthropic does not pick a winner. Its September 2025 post says many AI apps use "embedding-based pre-inference time retrieval," and that teams increasingly add "just in time" strategies on top. It calls Claude Code a hybrid.
The quote: "CLAUDE.md files are naively dropped into context up front, while primitives like glob and grep allow it to navigate its environment and retrieve files just-in-time, effectively bypassing the issues of stale indexing and complex syntax trees."
Stale indexing is the key cost of pre-built indexes. If the index was built an hour ago and you changed the file, the index is wrong. Grep reads the file as it is now.
The post also states the price of exploring: "runtime exploration is slower than retrieving pre-computed data." Without guidance an agent "can waste context by misusing tools, chasing dead-ends, or failing to identify key information." It adds that the hybrid approach "might be better suited for contexts with less dynamic content, such as legal or finance work." Its closing advice: "'do the simplest thing that works' will likely remain our best advice."
There is one more rule of thumb, from the 2024 post. If your knowledge base is smaller than 200,000 tokens (about 500 pages), "you can just include the entire knowledge base in the prompt," with no RAG needed. Context windows differ by model and have grown since, so check yours with the context window explainer.
What does Cursor say?
Cursor argues the other side, with a caveat. Its November 2025 blog post says its agent uses semantic search "in addition to the regex-based searching provided by a tool like grep." Semantic search matches by meaning, so it can answer "where do we handle authentication?" without knowing any names.
Its own test reports "on average 12.5% higher accuracy in answering questions (6.5%-23.5% depending on the model)." In an A/B test, agent-written code stayed in the codebase 0.3% more often with semantic search, and 2.6% more on codebases of 1,000 files or more. Cursor concludes: "Our agent makes heavy use of grep as well as semantic search, and the combination of these two leads to the best outcomes."
Cursor's own evaluation, 6 Nov 2025. Vendor-run, not independent. · aliteq research
Two cautions. These are Cursor's numbers about Cursor's agent. And the Cursor search docs page I read on 3 October 2026 describes an exact-match tool called Instant Grep and says "Cursor does not upload file paths or code to build a search index, and it does not store embeddings of your codebase for search." That page does not mention semantic search. I cannot tell from the docs whether the two statements describe different features, so I am not claiming Cursor dropped or kept it. Check its current docs before you decide anything about privacy.
Why does plain RAG still miss things?
Plain RAG can retrieve the wrong chunks. Anthropic found embeddings "can miss crucial exact matches," and gave an example: a search for "Error code TS-999" might return general error-code content and miss that exact string. It suggests adding BM25, an older keyword-matching method, alongside embeddings.
The post measured how often the right chunks failed to appear in the top 20 results. Embeddings alone missed 5.7%. Adding context to each chunk cut that to 3.7%, adding BM25 got 2.9%, and adding a reranking step got 1.9%.
Anthropic's own 2024 datasets, read 3 Oct 2026. Lower is better. · aliteq research
Notice what that means. The best RAG setup in that test still missed 1.9%. And the exact-match fix is, in effect, a keyword search, which is what grep is. The two approaches borrow from each other.
Which should you use? Five decision rules
Use grep and file reads when you can name the thing you are looking for. Use embeddings when you only know what it means. These rules come from the vendor guidance above, not from our own tests.
A function, error message, variable or ID is an exact string. Cursor's docs say 'the fastest way to find code is an exact match.'
'Where do we handle login?' or 'which policy covers refunds?' needs matching by meaning, which is what Cursor's semantic search targets.
Anthropic's 2024 rule of thumb was to put everything in the prompt under 200,000 tokens. Check your model's window.
Anthropic names stale indexing as a problem grep avoids. If you index, plan how and when it refreshes.
Anthropic found embeddings plus BM25 beat embeddings alone. Use both, and test on your own questions.
Agentic search vs RAG at a glance
How it finds things
Agentic search
Tools at runtime: grep, glob, file reads
RAG with a vector database
Embeddings matched by meaning
Setup
Agentic search
None beyond the tools
RAG with a vector database
Chunking, an embedding model, a vector database
Freshness
Agentic search
Reads the live file
RAG with a vector database
Index can go stale
Weak spot
Agentic search
Slower; can waste context chasing dead ends (Anthropic)
RAG with a vector database
Can miss exact matches (Anthropic)
Good fit
Agentic search
Code, logs, data you can name
RAG with a vector database
Large piles of docs; meaning-based questions
Agentic search
RAG with a vector database
How it finds things
Tools at runtime: grep, glob, file reads
Embeddings matched by meaning
Setup
None beyond the tools
Chunking, an embedding model, a vector database
Freshness
Reads the live file
Index can go stale
Weak spot
Slower; can waste context chasing dead ends (Anthropic)
Can miss exact matches (Anthropic)
Good fit
Code, logs, data you can name
Large piles of docs; meaning-based questions
Our summary of Anthropic and Cursor guidance, read 3 Oct 2026. Not a benchmark. · aliteq research
What does this mean if you are building with an AI tool?
If your AI coding tool is working on your own app, it already searches for you. You do not need to add a vector database so that Claude Code or Cursor can read your project. Their docs describe search as a built-in tool.
You would add RAG when your finished app needs to answer questions from a pile of content, such as help articles or uploaded PDFs. That is a product feature, not a development tool. Start small: if the content fits in the prompt, try that first. If you do need retrieval, a Postgres extension such as pgvector keeps it in a database you may already run.
Keep your standing instructions short and in a file the agent reads up front, as Anthropic does with CLAUDE.md. Our Claude Code guide for vibe coders shows how, and AI agents explained covers the loop of tools an agent runs.
Quick answers
What is agentic search?
The AI agent searches for itself using tools such as grep, file listing and file reads, then loads what it finds into context. Anthropic calls this a "just in time" approach. Nothing is indexed ahead of time.
Does Claude Code use RAG or a vector database?
Anthropic's September 2025 post describes Claude Code as a hybrid. It drops CLAUDE.md files into context up front and uses glob and grep to retrieve files on demand. The post does not describe a vector database for code search.
Is RAG dead?
No. Anthropic's own posts describe RAG as the typical answer for knowledge bases too large for a prompt, and Cursor reports gains from semantic search. The shift is toward combining RAG with agent-driven search, not abandoning either.
When do I need a vector database?
When people ask questions by meaning over a large pile of text you cannot fit in the prompt or explore by name. A small set of documents may fit in the prompt directly. Anthropic's 2024 threshold was 200,000 tokens, which depends on the model.
Why not use embeddings for everything?
Embeddings can miss exact matches such as an error code, and an index can go stale when files change. Anthropic found adding BM25 keyword matching to embeddings reduced misses, and it recommends the simplest approach that works.
Can I trust the 12.5% number from Cursor?
Treat it as a vendor result. Cursor ran the test on its own agent and published the numbers. It shows semantic search helped there, not that it will help on your project.