Skip to content
All articles
  • Applied AI
  • RAG
  • Document retrieval

PageIndex and the latest funeral for RAG

What document trees can do better than basic vector search, where the 98.7% claim comes from, and why free source code still comes with an inference bill.

Eduard Smirnov8 min read

The entire RAG industry is about to get cooked, apparently. No vector database. No embeddings. No chunking. No similarity search. A 98.7% benchmark score. Free and open source.

Somewhere, a perfectly functional search system is being dragged into a migration meeting.

PageIndex is worth looking at. It gives language models a structured way to navigate documents and retrieve evidence. The funeral announcement becomes awkward when you visit the repository and read how its own authors describe it: reasoning-based RAG.

RAG means retrieval-augmented generation. You retrieve information and use it to produce an answer. Changing how retrieval works still leaves you with RAG. There is no requirement to pay a vector database company before the acronym applies.

What PageIndex challenges is a familiar implementation: split a document into chunks, embed them, retrieve the nearest matches and ask a model to answer from whatever comes back. That implementation has weaknesses, especially with long reports. It also has more capable relatives than the five-chunk demo that keeps getting wheeled out for these comparisons.

Give the model a document it can navigate

A report already contains clues about how to read it. Headings, subsections, tables, footnotes and references to other pages tell you where information belongs and what it depends on. Flattening everything into isolated passages can lose those relationships.

PageIndex represents a document as a tree. The model examines the index, selects promising sections, retrieves their contents and keeps looking when the evidence is insufficient. Its technical introduction describes following internal references and fetching neighbouring sections when more context is needed.

The useful part is being able to act on something like "see Note 12". A reader understands that this is an instruction to look elsewhere. A basic similarity search may return the sentence and leave the actual note behind.

There is a reasonable engineering idea here. Calling it "reads like a human" is a convenient explanation of the navigation pattern. It does not mean the model has acquired the judgement of an accountant.

Revenue went up. Margin went down. Find the missing note

Consider a fictional manufacturer, Northbridge Components. All company details, figures, page numbers and retrieval paths in this example are invented. This is a worked example of the problem, not a PageIndex benchmark result.

An analyst asks:

Why did operating margin fall when revenue increased?

The evidence is spread across the annual report:

Location What it contains
Page 24, business review Revenue increased from GBP 1.6 billion to GBP 2 billion.
Page 81, income statement Operating profit increased from GBP 240 million to GBP 260 million.
Page 146, acquisition note Current-year operating costs include a GBP 38 million acquisition-related inventory charge.

The arithmetic is straightforward. The earlier margin was 240 / 1,600 = 15%. The current margin is 260 / 2,000 = 13%.

A basic retriever could find the revenue discussion and a paragraph about rising costs, then miss the acquisition note. The generator has enough material to produce something respectable about "integration pressures". It lacks the number that would make the explanation useful.

Illustrative evidence trail through a fictional annual report: revenue on page 24, operating profit on page 81, and a GBP 38 million acquisition charge on page 146. Together they explain the margin calculation and a major contributing cost.
An illustrative route through the fictional Northbridge report. The page numbers and figures are invented. Open the diagram to view it at full size.

A PageIndex workflow could navigate from the business review to the financial statements and then into the acquisition notes. With that evidence, it could explain that the reported margin fell by two percentage points and that the GBP 38 million charge represents 1.9% of current-year revenue.

There is a limit to that explanation. Adding the charge back gives an illustrative margin of 14.9%, but calling that the company's official adjusted margin would require evidence that the company uses that definition. Finding a footnote does not authorise inventing an accounting policy.

This is the advantage worth testing: the model can follow a document's structure to assemble connected evidence. It can also take a wrong turn. The example shows why the approach might help, not that the software must answer correctly.

The same problem turns up in a maintenance manual

Here is another hypothetical case. A technician asks why a pump shuts down after a firmware update even though the displayed temperature is normal.

The troubleshooting chapter says to check for overheating. A release note says the update changed which sensor controls shutdown. An appendix explains that the dashboard still displays the original sensor.

A weak retriever could return several convincing passages about cooling and ventilation. A workflow that follows the reference into the release notes and appendix could uncover the mismatch. Replacing a fan would be a fairly expensive way to discover that there were two temperature readings.

Again, vector search is not mathematically forbidden from finding those pages. A good system can preserve section context, combine keyword and vector search, and rerank candidates. Anthropic's contextual retrieval approach adds explanatory context to chunks and combines embeddings, BM25 keyword search and reranking.

That is a more useful comparison than an agent allowed to investigate a document versus a retriever that gets one attempt and five snippets. Keep the answer model and evaluation questions comparable, and report the extra search cost.

Beating the tutorial version of RAG is a useful result. It does not settle what to deploy.

What the 98.7% actually refers to

Vectify reports 98.7% accuracy for Mafin 2.5, a financial question-answering system built on PageIndex. The result belongs to that complete system and its evaluation setup. Installing the open-source package does not reproduce it. The distinction is explicit in the team's benchmark announcement.

The evaluation repository says it tests on the FinanceBench public set. The FinanceBench release contains 150 public examples; the broader collection contains 10,231 questions. "Full benchmark" in the vendor comparison needs to be read in the context of the public set being evaluated.

Context for the reported benchmark: 98.7 percent answer accuracy for Mafin 2.5, a system built on PageIndex, evaluated on the 150-example FinanceBench public set. The broader FinanceBench collection has 10,231 questions.
The reported score and its scope, from the Mafin evaluation and FinanceBench documentation. This is an explanatory graphic, not a new experiment.

The published grading script uses language models to judge answer equivalence. The repository also describes human annotations for ambiguous or disputed cases. Those choices are part of the measurement and deserve to travel with the number.

Publishing answers and grading code helps people inspect the claim. However, the comparison table includes systems evaluated on different proportions of the public set. It does not establish that PageIndex beats every vector RAG implementation under identical conditions.

The number fits beautifully in a social post. Apparently the experimental conditions exceeded the character limit.

The tree has its own failure modes

The model decides which branch to inspect. If a summary leaves out an exception, or a heading gives an unhelpful description, the model can skip the relevant section. A clear navigation trail makes that decision easier to investigate. It does not make the decision correct. This is a risk implied by the design, not a measured PageIndex failure rate.

There is also a practical difference between navigating one well-structured annual report and selecting the right documents from a large, messy collection. PageIndex has a separate file-system layer for searching across documents. That makes corpus selection another part of the system to evaluate. A convincing answer from the wrong year's report is still wrong.

The "no chunking" claim deserves a little care too. The system still has smaller addressable units, including sections and pages. Preserving meaningful boundaries is useful. A tree does not remove the need to decide which content belongs together or how much to send to the model.

Then there is the bill. Removing embeddings and a vector database removes particular expenses. Iterative model calls introduce others. Response time depends on the model, the document and how much searching a question requires. Measure that on the questions people actually ask.

The repository is MIT-licensed. It also documents model costs and requires an LLM connection for the full reasoning workflow. "Free source code" is accurate. "100% free" needs a conversation with whoever pays for inference or owns the machine running it.

A small test before the migration meeting

The current repository includes PageIndex Flash, which can extract a raw document tree without an LLM. Summaries and model-driven refinement are additional steps. The Flash documentation describes the distinction.

For a first inspection, use a public report or a document you are authorised to process:

git clone https://github.com/VectifyAI/PageIndex.git
cd PageIndex
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .

python run_pageindex.py \
  --mode flash \
  --pdf_path examples/documents/q1-fy25-earnings.pdf \
  --no-summary \
  --optimize off

That writes results/q1-fy25-earnings_structure.json. Inspect the titles and page ranges. Do they represent the document sensibly? A successfully generated JSON file establishes that indexing ran. It says nothing yet about answer accuracy.

For the full retrieval experiment, configure a supported language model using the repository's SDK instructions. Keep the raw tree test separate from the evaluation of summaries, search and answers, so it is clear what you have actually measured.

Use questions that require a buried exception, a calculation across sections, or evidence the report does not contain. Compare PageIndex with a competent hybrid-search baseline. Check whether each answer found all necessary evidence and whether its citations support the claims. Record response time and cost as well. A system that confidently answers the unanswerable question should lose points, however attractive the paragraph is.

The phrase "RAG is dead" keeps coming back because RAG has become shorthand for one particular recipe. Change the recipe and somebody declares the whole category obsolete. It also gets more attention than "we have a retrieval approach that deserves testing on long documents".

PageIndex gives a good reason to examine how much document structure your pipeline discards. For reports and manuals with useful internal references, that could matter a lot. For other workloads, the extra model-driven navigation may not earn its cost.

If the answer depends on a footnote on page 146, make sure the system reads page 146. The database obituary can wait.