aliteq.

Claude Code Doesn't Need a Vector Database. Does Your AI App? Agentic Search vs RAG

Your coding agent finds code by grepping and reading files, not by embeddings. Anthropic says that works. Cursor says semantic search added accuracy. Here is how to decide which your own app needs.

SyntaxUpdated 2d ago7 min readWeb story
Hand-drawn editorial illustration of a magnifying glass over a scatter of lavender dots with one coral dot connected by a line to a stack of paper files, on a near-black background
Share

You are using Cursor, Claude Code or Lovable, and you read that real AI apps need a vector database. Then you notice your coding agent seems to find code just fine without one. Which is it?

Both are true, for different jobs. This page explains the two approaches in plain English, shows what the two biggest vendors say, and ends with rules you can apply to your own app. New to the vocabulary? Start with what a context window is and context engineering for vibe coders.

Agentic search means the AI agent searches for itself, with tools, while it works. It runs grep to find a word, lists files by name, opens one, reads it, then decides what to look at next. Nothing is indexed ahead of time.

Anthropic calls the pattern "just in time" context. Its September 2025 engineering post says such agents "maintain lightweight identifiers (file paths, stored queries, web links, etc.) and use these references to dynamically load data into context at runtime using tools."

Grep is a command that finds lines matching a pattern. In Claude Code, the docs list a Search tool category that can "find files by pattern, search content with regex, explore codebases." On macOS, Linux and WSL, the tools reference says Claude searches with find and grep through the Bash tool.

What is RAG, and where does the vector database come in?

RAG stands for retrieval-augmented generation. You prepare your content ahead of time, then fetch the relevant pieces and paste them into the prompt before the AI answers. A vector database is where the prepared pieces live.

Anthropic's Contextual Retrieval post lists the steps. Break the content into chunks of "usually no more than a few hundred tokens." Use an embedding model to turn each chunk into a vector, a list of numbers that encodes meaning. Store them in a vector database. At question time, find the chunks closest in meaning to the question and add them to the prompt.

You can try this in a database you may already use. Supabase's docs say vectors there are enabled via pgvector, "a Postgres extension for storing and querying vectors." One more fact: the Claude API docs say "Anthropic does not offer its own embedding model," and point to Voyage AI as one provider. For a hands-on version, see how to chat with your documents locally.

What does Anthropic say about when to use each?

Anthropic does not pick a winner. Its September 2025 post says many AI apps use "embedding-based pre-inference time retrieval," and that teams increasingly add "just in time" strategies on top. It calls Claude Code a hybrid.

The quote: "CLAUDE.md files are naively dropped into context up front, while primitives like glob and grep allow it to navigate its environment and retrieve files just-in-time, effectively bypassing the issues of stale indexing and complex syntax trees."

Stale indexing is the key cost of pre-built indexes. If the index was built an hour ago and you changed the file, the index is wrong. Grep reads the file as it is now.

The post also states the price of exploring: "runtime exploration is slower than retrieving pre-computed data." Without guidance an agent "can waste context by misusing tools, chasing dead-ends, or failing to identify key information." It adds that the hybrid approach "might be better suited for contexts with less dynamic content, such as legal or finance work." Its closing advice: "'do the simplest thing that works' will likely remain our best advice."

There is one more rule of thumb, from the 2024 post. If your knowledge base is smaller than 200,000 tokens (about 500 pages), "you can just include the entire knowledge base in the prompt," with no RAG needed. Context windows differ by model and have grown since, so check yours with the context window explainer.

What does Cursor say?

Cursor argues the other side, with a caveat. Its November 2025 blog post says its agent uses semantic search "in addition to the regex-based searching provided by a tool like grep." Semantic search matches by meaning, so it can answer "where do we handle authentication?" without knowing any names.

Its own test reports "on average 12.5% higher accuracy in answering questions (6.5%-23.5% depending on the model)." In an A/B test, agent-written code stayed in the codebase 0.3% more often with semantic search, and 2.6% more on codebases of 1,000 files or more. Cursor concludes: "Our agent makes heavy use of grep as well as semantic search, and the combination of these two leads to the best outcomes."

Three stat cards from Cursor's own evaluation: 12.5 percent average accuracy gain with semantic search, 0.3 percent more agent-written code retained overall, and 2.6 percent more on codebases of 1,000 or more files.
Cursor's own evaluation, 6 Nov 2025. Vendor-run, not independent. · aliteq research

Two cautions. These are Cursor's numbers about Cursor's agent. And the Cursor search docs page I read on 3 October 2026 describes an exact-match tool called Instant Grep and says "Cursor does not upload file paths or code to build a search index, and it does not store embeddings of your codebase for search." That page does not mention semantic search. I cannot tell from the docs whether the two statements describe different features, so I am not claiming Cursor dropped or kept it. Check its current docs before you decide anything about privacy.

Why does plain RAG still miss things?

Plain RAG can retrieve the wrong chunks. Anthropic found embeddings "can miss crucial exact matches," and gave an example: a search for "Error code TS-999" might return general error-code content and miss that exact string. It suggests adding BM25, an older keyword-matching method, alongside embeddings.

The post measured how often the right chunks failed to appear in the top 20 results. Embeddings alone missed 5.7%. Adding context to each chunk cut that to 3.7%, adding BM25 got 2.9%, and adding a reranking step got 1.9%.

Bar chart of top-20 retrieval failure rate in Anthropic's 2024 test: embeddings only 5.7 percent, plus contextual chunks 3.7, plus BM25 2.9, plus reranking 1.9.
Anthropic's own 2024 datasets, read 3 Oct 2026. Lower is better. · aliteq research

Notice what that means. The best RAG setup in that test still missed 1.9%. And the exact-match fix is, in effect, a keyword search, which is what grep is. The two approaches borrow from each other.

Which should you use? Five decision rules

Use grep and file reads when you can name the thing you are looking for. Use embeddings when you only know what it means. These rules come from the vendor guidance above, not from our own tests.

A function, error message, variable or ID is an exact string. Cursor's docs say 'the fastest way to find code is an exact match.'

'Where do we handle login?' or 'which policy covers refunds?' needs matching by meaning, which is what Cursor's semantic search targets.

Anthropic's 2024 rule of thumb was to put everything in the prompt under 200,000 tokens. Check your model's window.

Anthropic names stale indexing as a problem grep avoids. If you index, plan how and when it refreshes.

Anthropic found embeddings plus BM25 beat embeddings alone. Use both, and test on your own questions.

Agentic search vs RAG at a glance

How it finds things

Agentic search
Tools at runtime: grep, glob, file reads
RAG with a vector database
Embeddings matched by meaning

Setup

Agentic search
None beyond the tools
RAG with a vector database
Chunking, an embedding model, a vector database

Freshness

Agentic search
Reads the live file
RAG with a vector database
Index can go stale

Weak spot

Agentic search
Slower; can waste context chasing dead ends (Anthropic)
RAG with a vector database
Can miss exact matches (Anthropic)

Good fit

Agentic search
Code, logs, data you can name
RAG with a vector database
Large piles of docs; meaning-based questions
Scorecard comparing grep with embeddings: grep is the best fit for exact strings and always reads the live file with no setup; embeddings are the best fit for meaning-based questions and large document piles but need a chunker, embedder and database and can go stale.
Our summary of Anthropic and Cursor guidance, read 3 Oct 2026. Not a benchmark. · aliteq research

What does this mean if you are building with an AI tool?

If your AI coding tool is working on your own app, it already searches for you. You do not need to add a vector database so that Claude Code or Cursor can read your project. Their docs describe search as a built-in tool.

You would add RAG when your finished app needs to answer questions from a pile of content, such as help articles or uploaded PDFs. That is a product feature, not a development tool. Start small: if the content fits in the prompt, try that first. If you do need retrieval, a Postgres extension such as pgvector keeps it in a database you may already run.

Keep your standing instructions short and in a file the agent reads up front, as Anthropic does with CLAUDE.md. Our Claude Code guide for vibe coders shows how, and AI agents explained covers the loop of tools an agent runs.

Quick answers

What is agentic search?
The AI agent searches for itself using tools such as grep, file listing and file reads, then loads what it finds into context. Anthropic calls this a "just in time" approach. Nothing is indexed ahead of time.
Does Claude Code use RAG or a vector database?
Anthropic's September 2025 post describes Claude Code as a hybrid. It drops CLAUDE.md files into context up front and uses glob and grep to retrieve files on demand. The post does not describe a vector database for code search.
Is RAG dead?
No. Anthropic's own posts describe RAG as the typical answer for knowledge bases too large for a prompt, and Cursor reports gains from semantic search. The shift is toward combining RAG with agent-driven search, not abandoning either.
When do I need a vector database?
When people ask questions by meaning over a large pile of text you cannot fit in the prompt or explore by name. A small set of documents may fit in the prompt directly. Anthropic's 2024 threshold was 200,000 tokens, which depends on the model.
Why not use embeddings for everything?
Embeddings can miss exact matches such as an error code, and an index can go stale when files change. Anthropic found adding BM25 keyword matching to embeddings reduced misses, and it recommends the simplest approach that works.
Can I trust the 12.5% number from Cursor?
Treat it as a vendor result. Cursor ran the test on its own agent and published the numbers. It shows semantic search helped there, not that it will help on your project.

Found this useful? Share it

Share
Syntax

Build Editor

Syntax

I explain what's actually happening when you build software by talking to an AI — what the model is doing, what's really running your app, and where the sharp edges are. No jargon without a picture, no hype, and an honest 'hire someone' when that's the answer.

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading