August 15, 2026 · By YasKad
VectifyAI/PageIndex

PageIndex: document RAG that traverses trees instead of vectors

VectifyAI/PageIndex · 35,848★ · 3,158 forks

Everything worth knowing about VectifyAI/PageIndex: a hierarchical index for retrieving evidence from long documents through language-model reasoning, without a vector database as a requirement.


What PageIndex is

PageIndex is Vectify AI’s open-source component for building a tree-shaped index from a PDF or Markdown file and using it in retrieval-augmented generation (RAG). The tree resembles an enriched index or table of contents: nodes carry a title, page range, summary, and identifier. The model can navigate that structure and retrieve specific sections instead of comparing the query against chunk vectors.

The pitch does not mean no model is involved: parsing and generating the structure use a language model. “Without vectors” means the published pipeline does not require embeddings, similarity search, or a vector database. The README places the use case in financial reports, legal and regulatory documents, manuals, medical literature, and long technical books.

There are three related offerings: the self-hostable repository, a cloud service with enhanced OCR and processing, and PageIndex Chat. The official site also presents MCP and API integration; enterprise options include a dedicated cluster or private VPC deployment.

The origin: a critique that similarity is not relevance

The repository was created on April 1, 2025. The README attributes the technical framing to Mingtian Zhang, Yu Tang, and the PageIndex team, and links to a project introduction dated September 2025. Those names are therefore documented as authors of the presentation, not necessarily the entire founding team.

The launch narrative starts from a specific tension with vector-based RAG: Vectify AI argues that semantic similarity does not guarantee contextual relevance in long professional documents. Its declared inspiration is the way an expert reader works: first build a table of contents, then reason over its branches.

The Hacker News discussion shows that thesis drew interest and friction from the start. In the August 2025 launch thread, the author replied that initial tree construction can be slower than the vector approach and that a large tree can slow down retrieval; the design prioritizes precision over speed. That is a limitation the project itself acknowledges, not a promise to universally replace vector databases.

Digital scale comparing a glowing hierarchical tree, representing precision and structure, with a fast-moving vector cloud, representing speed and similarity; the scale tips slightly toward the tree.

Philosophy and principles

The verifiable principles from the README and documentation can be summarized as follows:

  • Structure over artificial fragmentation: preserve natural sections and their hierarchy, instead of splitting text purely by length.

Conceptual representation of "structure over artificial fragmentation": a continuous scroll of glowing digital text organically folds into a neon-outlined tree hierarchy instead of being chopped into arbitrary chunks.

  • Reasoning over similarity: traverse a hierarchical index to decide which branches contain relevant evidence.

Conceptual illustration of "reasoning over similarity": on the left, a glowing AI brain navigates a neon tree index; on the right, a scattered cluster of vector dots connected by chaotic lines fades into the darkness.

  • Traceable evidence: return explicit references to sections and pages so retrieval can be inspected.

Close-up of a glowing node in the PageIndex tree projecting a holographic card with a page number, section title, and a brief summary, representing traceable evidence.

  • Conversation context: selection can incorporate conversational history and domain knowledge, according to the project’s documentation.
  • Choice based on the operational trade-off: the project does not claim the tree is always faster. The author stated on HN that a small tree can be efficient, while for speed over precision he recommends a vector database.

The result is a design alternative, not independent proof that any vector-based RAG is opaque or inferior. The 98.7% figure on FinanceBench comes from the project itself and its VectifyAI/Mafin2.5-FinanceBench repository; it should be read as a result published by Vectify AI, not as general external validation.

How it works

The documented local workflow has two phases:

  1. Index generation: PageIndex processes the document and produces a semantic tree. For a PDF, the output file represents the hierarchy, its summaries, and page ranges. It can detect an existing table of contents or build the structure from the document itself.
  2. Reasoned retrieval: given a question, a model examines the tree and follows branches until it locates the relevant sections. The justification is tied to pages and nodes, rather than being limited to a similarity score.

Split-screen image: on the left, a PDF turns into a semantic tree; on the right, an AI agent queries that tree, following a glowing path down the branches to a specific node.

The main program is run_pageindex.py. It accepts PDFs via --pdf_path and Markdown via --md_path; for Markdown, the #, ##, and further heading levels are interpreted as the hierarchy. The PageIndex Flash preview is enabled with --flash: it uses heuristics to extract the structure and reserves the model for summaries; --optimize adds a refinement pass.

The repository also includes a self-hosted agentic RAG example built with the OpenAI Agents SDK at examples/agentic_vectorless_rag_demo.py, plus notebooks for vectorless RAG and visual RAG. The latter works over page images, according to the README.

Representation of an agentic workflow without a vector database: a glowing robotic hand points at the branches of a holographic table-of-contents tree and pulls out glowing page panels.

Official and semi-official status

PageIndex is an official Vectify AI project: the repository sits under the VectifyAI organization, the site and docs present it as a Vectify AI product, and the README links to its own Chat, MCP, and API services.

No evidence was recovered that PageIndex has been accepted into an official marketplace run by a model provider. MCP is an open integration standard, not an endorsement from Anthropic, OpenAI, or any other provider. In practice, the official VectifyAI/pageindex-mcp server and the official guides allow it to be connected to agents that support MCP; that establishes documented compatibility, not a performance certification or de facto standard status.

The ecosystem

Vectify AI repositories

Querying the organization’s API identified the following related public repositories; star counts are from this run’s measurement:

  • VectifyAI/pageindex-mcp — PageIndex’s MCP server for retrieval via tree search and reasoning; 377 stars.
  • VectifyAI/pageindex-js-sdk — TypeScript SDK for PageIndex document processing; 7 stars.
  • VectifyAI/OpenKB — a knowledge base for language models that compiles documents into an interconnected wiki; 3,261 stars.
  • VectifyAI/ChatIndex — tree indexing and retrieval for long conversational memory; 156 stars.
  • VectifyAI/ConDB — a KV-cache-native context database for tree-based retrieval; 51 stars.
  • VectifyAI/Mafin2.5-FinanceBench — the FinanceBench evaluation for Mafin 2.5, listed by the organization as powered by PageIndex; 91 stars.
  • VectifyAI/MMLongBench-Doc-V2 — a document evaluation repository with no stars in the queried API response.

Abstract representation of the PageIndex ecosystem: a central glowing tree connected via data streams to chat, cloud server, and API gateway icons, referencing MCP integration.

The README explicitly lists OpenKB, ChatIndex, ConDB, and PageIndex MCP as long-context infrastructure. It does not claim they are mandatory dependencies for running run_pageindex.py.

Derivatives, ports, and translations

A repository name search recovered VectifyAI/pageindex-mcp as the directly identifiable derivative. osdotsystem/pageindex-open also appeared in a February 2026 HN submission as an open implementation; its README and GitHub metadata were not recovered in this run, so its technical relationship, maintenance, and metrics remain unverified. No translations or ports in languages other than English were identified with sufficient evidence to name them as such.

Repo numbers

Measurement: August 5, 2026; public GitHub API.

MetricValue
Stars35,025
Forks3,070
Real subscribers141
Commits357
Open issues reported by the API146
Primary languagePython
LicenseMIT
CreatedApril 1, 2025
Last code pushAugust 4, 2026
Last metadata updateAugust 5, 2026
Latest GitHub releasev0.3.0.dev3, July 10, 2026
Latest PyPI version0.2.8

The top contributors returned by the API were rejojer (233 contributions), zmtomorrow (68), BukeLy (23), and S3DFX-CYBER (16). The 357-commit total was obtained from the last-page link of GitHub’s commit pagination. open_issues_count can include open pull requests. Likewise, watchers_count duplicates the star count here, which is why the table reports the independent subscribers_count field as the real subscriber figure.

How to contribute

There is no CONTRIBUTING.md or contribution guide in the recovered root directory, so an official fork, branch, or pull-request-template workflow cannot be attributed. There are tests/ and .github/ directories, and the visible pull requests show technical contributions.

For example, pull request #188 fixed a KeyError and context exhaustion while processing documents of roughly 800 pages; it added 14 mocked tests and was merged on July 3, 2026. This shows the repository uses tests for significant changes, but it does not substitute for a published contribution policy.

How the community received it

The recovered reception is substantial and, above all, concrete on Hacker News:

  • Submission 45036944, posted by page_index on August 27, 2025, reached 192 points and 128 comments. brap found the navigation process natural and considered the indexing and search cost acceptable for their case. agentcoops valued that, for internal RAG, one can buy a predictably better answer with more compute; that is an individual opinion, not a benchmark.
  • In the same thread, marcodena objected that vectors are far more efficient. mosselman asked about latency and cost under heavy search volume, and suggested alternatives based on pre-indexing and BM25. neya argued the approach may fit background queues but not near-instant on-demand RAG, given the trade-off between speed, precision, and cost.
  • ineedasername questioned using the “vectorless” label as a synonym for a lesser approach: after reviewing the code, they argued the core value is generating and traversing a table of contents through iterative model calls. gillesjacobs asked for comparisons against well-tuned hybrid retrievers and noted that the Mafin 2.5 evaluation is narrow. These are technical criticisms attributed to their authors; no independent evaluation resolving them was recovered.
  • On Reddit, the thread r/Rag: 1n1iqy3, posted by CathyCCCAAAI on August 27, 2025, logged 127 points and 57 comments. The opening post praises transparency and traceability, but it is promotional material, not an independent review. Another result, r/LovingOpenSourceAI: 1t4il64, reached 52 points and 41 comments; no opinions from its comments are attributed here because they were not individually recovered.

Several lower-reach HN submissions tied to PageIndex MCP were also found. The most visible, 45482992, had 14 points and 3 comments. That is a signal of the MCP add-on’s reach, not proof of mass adoption.

No verifiable posts on X were recovered that would allow attributing a specific reaction to an account, nor a Product Hunt launch page; the Product Hunt search returned no products for “PageIndex VectifyAI.” The automated YouTube search returned a technical page without recoverable titles or view counts, so no tutorials or view figures are invented here.

PageIndex versus other approaches

ApproachVerifiable relationshipVerifiable difference or limit
Vector-based RAG with a vector databaseRetrieves context for a language model and is the README’s explicit point of contrast.PageIndex structures a tree and uses tree-based reasoning; the vector approach retrieves by similarity. The author himself notes the vector option may be preferable when speed is the priority.
GraphRAGagentcoops compared it in thread 45036944 as a RAG alternative with preprocessing.The cost and scalability comparison is the commenter’s opinion, not a recovered benchmark.
neuml/txtaiIts author, dmezzetti, named it in the thread as a lightweight retrieval option that does not require a separate vector database.The source does not support concluding functional equivalence or measuring performance against PageIndex.
datalab-to/markerleetharris recommended it in the discussion for self-hosted document extraction.It refers to document extraction, not a direct competitor for reasoned retrieval; it should not be confused with a full RAG system.

Quick usage guide

Installation and first local run

The README publishes the following minimal workflow:

git clone https://github.com/VectifyAI/PageIndex.git
cd PageIndex
pip3 install --upgrade -r requirements.txt

Create .env at the root and add a model key compatible with LiteLLM; the official example uses:

OPENAI_API_KEY=your_openai_key_here

Then run:

python3 run_pageindex.py --pdf_path /path/to/document.pdf

The first expected result is a tree structure of the document. For Markdown, the equivalent official command is:

python3 run_pageindex.py --md_path /path/to/document.md

Common workflows

  1. Index a PDF report: run python3 run_pageindex.py --pdf_path /path/report.pdf; use the generated tree to inspect titles, nodes, summaries, and page ranges.

  2. Speed up structure extraction: run python3 run_pageindex.py --flash --pdf_path /path/report.pdf. Add --optimize when you want an additional refinement pass with the model.

  3. Run the self-hosted agentic example: install the optional dependency and start the official example:

    pip3 install openai-agents
    python3 examples/agentic_vectorless_rag_demo.py
  4. Integrate the hosted service: create a key from the developer dashboard, upload and process the document, and choose MCP so your own agent can invoke PageIndex as a tool, or the Chat API to receive answers from PageIndex’s hosted model.

Essential configuration

  • .env: stores the model provider key for the local run.
  • --model: selects the model; the documented default is gpt-4o-2024-11-20.
  • --toc-check-pages: limits how many pages are inspected when looking for a table of contents; default: 20.
  • --max-pages-per-node: limits pages per node; default: 10.
  • --max-tokens-per-node: limits tokens per node; default: 20,000.
  • --if-add-node-id, --if-add-node-summary, and --if-add-doc-description: toggle identifiers, summaries, and document description; all default to yes.

Common pitfalls and fixes

  • Markdown auto-converted from PDF or HTML: the README advises against using --md_path if the conversion did not preserve the hierarchy. Since Markdown mode infers levels from #, a flawed structure produces a flawed tree. The documented alternative is to process the PDF directly or obtain Markdown that preserves the hierarchy.
  • Complex or scanned PDFs: the repository uses standard PDF parsing. Vectify AI states that its cloud service provides enhanced OCR, tree construction, and retrieval; that is the documented path when local extraction is not enough.
  • Very large documents: merged pull request #188 fixed context failures on documents of roughly 800 pages. Update to a revision that includes that fix and test with the size and model of your own environment.
  • Expecting universally low latency: HN records reasonable objections about cost and scale. The author himself states large trees can be slower; measure indexing and query cost against your own corpus before replacing an existing vector retriever.

Integrations and migration

The repository provides an example with the OpenAI Agents SDK; the official service offers MCP and an API. MCP lets compatible agents or frameworks, including those using LangChain or the OpenAI Agents SDK per the documentation, call PageIndex as a tool without a specific manual integration.

No official migration guide to or from LangChain, LlamaIndex, GraphRAG, or a vector database was recovered. The prudent migration path is to run both retrievers over the same set of queries and compare latency, cost, traceability, and accuracy on your own domain; that recommendation is operational analysis, not an official PageIndex procedure.

Use cases and who this repository can help

  • Finance, compliance, or legal teams working with reports, regulatory filings, and long contracts can produce a tree with page ranges and keep navigable evidence for their answers. The README explicitly cites financial reports and legal and regulatory documents.
  • Developers of document agents can connect the service to an agent via MCP or reuse the OpenAI Agents SDK example, making retrieval a tool within an agentic sequence.
  • Teams that need to review why a passage was retrieved can benefit from nodes, summaries, and page references instead of accepting a list of similarity-ranked fragments with no visible hierarchy.
  • Users working with simple PDFs or well-structured Markdown can try the local workflow with run_pageindex.py; those with complex scans or private-deployment requirements should evaluate the OCR service and the enterprise options Vectify AI documents.
  • Teams with high query volume or strict latency requirements should not assume PageIndex replaces a vector database: the recovered comments and the author’s own reply point to cost and latency as relevant trade-offs.

Resources


Note: this article combines the PageIndex README, documentation, and official site, the public GitHub API, PyPI, Hacker News, and Reddit results consulted on August 5, 2026. Figures and availability change over time.

Comments