~/log/vectorless-rag
Vectorless RAG: I stopped chunking my documents and my RAG got better
AI 7 min read
If you’ve built any Retrieval-Augmented Generation (RAG) system in the last two years, you know the recipe by heart: split your document into chunks, embed each chunk into a vector, throw the vectors into a vector database, and at query time find the chunks whose embeddings are closest to the question’s embedding. It works well enough for a lot of use cases, and it’s become the default — to the point that “RAG” and “vector search” are almost treated as synonyms.
But that recipe has a quiet flaw baked into it: it retrieves what is semantically similar to the question, not necessarily what is relevant to answering it. Similarity and relevance sound like they should be the same thing, and often they are, but not always. A question about “why did the company’s margins shrink” might be best answered by a paragraph that never uses the words “margin” or “shrink” at all, because it’s actually describing a supply chain disruption three pages earlier. A vector search built on wording similarity can miss that connection entirely, while a human reader flipping through the document would catch it immediately — because they understand the structure and the argument, not just the words.
This is the exact gap that PageIndex, an open-source project from Vectify AI, was built to close. I used it as the basis for a small hands-on project to actually understand the idea at a level deeper than reading a blog post, and this write-up is what I learned building it.
What PageIndex actually does
PageIndex’s core claim is captured in one line from its own documentation:
Similarity ≠ relevance and relevance requires reasoning.
Instead of chunking a document and embedding the chunks, PageIndex builds a hierarchical tree index of the document — essentially an LLM-generated table of contents, where every node has a title, a page range, and a short summary of what that section covers. Retrieval then becomes a two-step process:
- Build the tree — parse the document into sections and generate a summary for each node, so the whole document’s structure is captured in a compact, navigable form.
- Search the tree — at query time, an LLM reads the node titles and summaries (never the raw chunk text) and reasons about which branch of the tree is likely to contain the answer, the same way a person would flip to a specific chapter instead of scanning every page.
Once the LLM has narrowed down to the relevant section(s), only then does it read the full, unbroken text of those sections to actually generate an answer. No vector database, no embedding model, and importantly, no arbitrary chunk boundaries slicing a sentence or an argument in half.
The team reports this approach achieved 98.7% accuracy on FinanceBench, a benchmark built from real financial filings, notably outperforming vector-based RAG baselines on the same documents — exactly the kind of long, structurally dense, professional document where “which section is this in” matters more than “which sentence sounds similar.”
Building a vectorless RAG from scratch
Reading about an idea and implementing it are different levels of understanding, so instead of just running PageIndex’s own scripts, I built a minimal version of the pipeline myself in a Jupyter notebook, using an unusual test document: the first Harry Potter book. The goal wasn’t to stress-test enterprise-scale retrieval — it was to build every piece by hand so I actually understood what’s happening at each stage, and to run a head-to-head comparison against a classic vector RAG baseline on the exact same content.
1. The tree builder
Most well-formed PDFs carry an embedded outline (bookmarks) — chapter titles and their starting pages added by whoever produced the file. I used that outline as a shortcut for chapter boundaries. (PageIndex’s own open-source pipeline does this detection itself, via an LLM scanning the first N pages when there’s no embedded outline.) For each chapter, I asked an LLM to generate a two-to-three sentence summary focused on concrete plot events and named characters, since that summary is the only thing the retriever ever sees when deciding where to look.
nodes = []
for i, (title, start_page) in enumerate(raw_chapters):
text = "\n".join(page_texts[start_page:end_page + 1])
nodes.append({
"node_id": f"{i+1:04d}",
"title": title,
"text": text,
"summary": summarize_node(text),
}) 2. The tree-search retriever
This is the heart of vectorless RAG. Given a question, the LLM is shown only the chapter titles and their summaries — never the full text — and asked to reason about which chapter(s) are likely to contain the answer. Critically, I asked it to return its reasoning alongside its choice, in structured JSON:
{
"reasoning": "The question asks about Harry's birthday gift. Chapter 4's summary mentions Hagrid arriving on Harry's birthday and giving him a creature-related gift…",
"selected_node_ids": ["0004"]
} That reasoning trail is what makes the retrieval traceable and explainable — you can literally read why a section was chosen, instead of trusting an opaque similarity score.
3. Answer generation
Once the relevant chapter(s) are identified, their full, unmodified text is passed to the LLM to actually answer the question. Because retrieval happens at the section level rather than the chunk level, there’s no risk of a chunk boundary cutting off half of the relevant sentence — a subtle but real failure mode in vector RAG that’s easy to overlook until you see it happen.
4. The vector RAG baseline, for comparison
To make the comparison fair and concrete, I built a standard vector RAG pipeline on the exact same book: fixed-size chunks, a local sentence-transformer embedding model, and cosine-similarity retrieval of the top-k chunks for each question.
Vectorless vs. vector RAG, side by side
Running both pipelines against the same set of questions made the difference tangible rather than theoretical.
The most revealing test wasn’t even a hard question — it was a trick one. I asked both systems about Mrs. Norris being petrified, an event that happens in the second Harry Potter book, not the one I’d indexed. The vector RAG baseline still confidently returned the most similar-sounding chunk it could find and attempted an answer, because cosine similarity always returns something. It has no concept of “none of this is actually relevant.” The tree-search approach was far more likely to recognize that no chapter’s summary matched the premise of the question at all, because it was reasoning about what the book contains, not just pattern-matching wording.
On more ordinary questions — “What does Harry see in the Mirror of Erised?”, “What creature does Harry get as a birthday present?” — both approaches did reasonably well, which makes sense: in a single short novel, plenty of questions have direct, similarity-friendly wording overlap with the answer. The gap widens specifically on indirect or inferential questions, and on documents that are longer, denser, and more structurally organized — which is exactly the professional-document use case (financial filings, legal contracts, technical manuals) PageIndex is designed for.
What this project actually taught me
Retrieval is a reasoning problem, not just a search problem. The most useful realization wasn’t “vectors bad, trees good” — it was that retrieval quality is fundamentally bottlenecked by how well the retriever understands the question’s intent, and an LLM reading summaries and reasoning about relevance can capture intent in a way a fixed embedding space can’t.
Traceability is a real, practical advantage. Being able to print out “here’s why I picked this chapter” is genuinely useful for debugging a RAG pipeline. With vector search, a bad retrieval is a mystery you have to reverse-engineer from cosine scores.
Vectorless RAG isn’t a strict upgrade, it’s a different trade-off. Tree-search retrieval costs more LLM calls per query and depends on the tree being built well in the first place. For huge corpora where low-latency, high-volume search matters more than deep reasoning per query, vector search still earns its place.
No-chunking matters more than it sounds. Watching an answer come from a full, unbroken section instead of a 500-character window made it obvious how much silent damage arbitrary chunk boundaries can do to context — damage that vector RAG just quietly absorbs as “normal.”
Closing thought
The vector-database era of RAG made retrieval feel like a solved, mechanical problem: embed, store, search, done. PageIndex is a useful reminder that it never really was solved — it was approximated. Reasoning-based retrieval brings back some of the cost and complexity that made vector search appealing in the first place, but for the documents where getting the right section actually matters, that trade-off is starting to look worth it.
The code for this project is on GitHub: anshujod/vectorless-rag. This post also appears on Medium.
~/related