linkgo
PageIndex

PageIndex

AI

Vectorless, reasoning-based RAG engine that indexes long documents as a tree and lets an LLM reason through it, with traceable citations.

-(0 Reviews)
Free Available
Starting from Free
Premium plans available

About PageIndex

PageIndex is a retrieval engine for long, complex professional documents that throws out the vector database entirely. Its premise is that vector RAG retrieves by semantic similarity while what retrieval actually needs is relevance, and relevance requires reasoning, so similarity search misses passages that are relevant but not similar. Inspired by AlphaGo, PageIndex instead builds a hierarchical tree index that follows the document's own natural sections, then has an LLM agentically search that tree the way a human expert turns to the right part of a long report. There is no chunking, no embedding and no vector store, and every answer is traceable to explicit page-level references rather than opaque vibe retrieval. It reached a state-of-the-art 98.7% accuracy on FinanceBench, the financial document QA benchmark. The open-source Python SDK runs fully locally with your own model key at roughly $0.001 per page to index, while PageIndex Cloud adds production OCR, image understanding, managed storage, line-level citations, an MCP server and the PageIndex File System, a file-level tree layer that scales reasoning across millions of documents. Access comes in three shapes: a chat platform for non-developers, an API and MCP server for developers integrating with Claude, Cursor, the OpenAI Agents SDK or their own framework, and enterprise VPC or on-premises deployment.

Screenshots

PageIndex screenshot 1
+

Key Features

Tree Index Instead of Vectors: Builds a hierarchical index from the document's own sections, so there is no chunking, no embeddings and no vector database to maintain.
Reasoning-Based Retrieval: An LLM agentically searches the tree using full context including conversation history and domain knowledge, rather than matching a query embedding.
Traceable Citations: Answers carry explicit page-level references locally and line-level citations on Cloud, so every claim can be checked against the source.
PageIndex Flash: Extracts tree structure from PDFs in seconds using the document's own layout information instead of building it with an LLM.
Local or Cloud SDK: pip install pageindex runs indexing, retrieval and chat entirely on your machine with your own key, or points the same client at PageIndex Cloud with an API key.
MCP Server and API: Connect document reasoning to Claude, Claude Desktop, Cursor or any MCP client, with API-key auth for developers and OAuth for chat users.
PageIndex File System: A Cloud-only file-level tree indexing layer that lets retrieval reason across an entire corpus rather than one document at a time.
Agent Framework Integrations: Ships integration paths for the OpenAI Agents SDK, the Anthropic SDK tool runner, the Claude Agent SDK and other frameworks.

Use Cases

Financial Document QA: Answer questions about 10-Ks, earnings reports and filings with the page the figure came from, the workload where it set a 98.7% FinanceBench record.
Legal and Regulatory Review: Retrieve the governing clause from contracts and regulatory filings where the relevant section is rarely the most semantically similar one.
Technical Manual Lookup: Find the correct procedure in long technical manuals where context and document hierarchy determine which section actually applies.
Medical and Academic Research: Reason over medical literature and textbooks that exceed a model's file size limits, with verifiable references.
Agent Document Tooling: Give an AI agent long-document reasoning via MCP so it can handle PDFs that models cannot ingest directly.
Enterprise Knowledge Bases: Index large document collections in the cloud with OCR and image understanding, and reason across the whole corpus with the File System layer.

Frequently asked questions about PageIndex

What is PageIndex?

Vectorless, reasoning-based RAG engine that indexes long documents as a tree and lets an LLM reason through it, with traceable citations.

How does PageIndex work?

PageIndex works by combining Tree Index Instead of Vectors: Builds a hierarchical index from the document's own sections, so there is no chunking, no embeddings and no vector database to maintain., Reasoning-Based Retrieval: An LLM agentically searches the tree using full context including conversation history and domain knowledge, rather than matching a query embedding., Traceable Citations: Answers carry explicit page-level references locally and line-level citations on Cloud, so every claim can be checked against the source., PageIndex Flash: Extracts tree structure from PDFs in seconds using the document's own layout information instead of building it with an LLM., Local or Cloud SDK: pip install pageindex runs indexing, retrieval and chat entirely on your machine with your own key, or points the same client at PageIndex Cloud with an API key. to help users with Financial Document QA: Answer questions about 10-Ks, earnings reports and filings with the page the figure came from, the workload where it set a 98.7% FinanceBench record., Legal and Regulatory Review: Retrieve the governing clause from contracts and regulatory filings where the relevant section is rarely the most semantically similar one., Technical Manual Lookup: Find the correct procedure in long technical manuals where context and document hierarchy determine which section actually applies., Medical and Academic Research: Reason over medical literature and textbooks that exceed a model's file size limits, with verifiable references., Agent Document Tooling: Give an AI agent long-document reasoning via MCP so it can handle PDFs that models cannot ingest directly..

What are the main features of PageIndex?

Key features include Tree Index Instead of Vectors: Builds a hierarchical index from the document's own sections, so there is no chunking, no embeddings and no vector database to maintain., Reasoning-Based Retrieval: An LLM agentically searches the tree using full context including conversation history and domain knowledge, rather than matching a query embedding., Traceable Citations: Answers carry explicit page-level references locally and line-level citations on Cloud, so every claim can be checked against the source., PageIndex Flash: Extracts tree structure from PDFs in seconds using the document's own layout information instead of building it with an LLM., Local or Cloud SDK: pip install pageindex runs indexing, retrieval and chat entirely on your machine with your own key, or points the same client at PageIndex Cloud with an API key..

Who is PageIndex for?

PageIndex is useful for anyone interested in Financial Document QA: Answer questions about 10-Ks, earnings reports and filings with the page the figure came from, the workload where it set a 98.7% FinanceBench record., Legal and Regulatory Review: Retrieve the governing clause from contracts and regulatory filings where the relevant section is rarely the most semantically similar one., Technical Manual Lookup: Find the correct procedure in long technical manuals where context and document hierarchy determine which section actually applies., Medical and Academic Research: Reason over medical literature and textbooks that exceed a model's file size limits, with verifiable references., Agent Document Tooling: Give an AI agent long-document reasoning via MCP so it can handle PDFs that models cannot ingest directly..

How much does PageIndex cost?

PageIndex offers a free tier with paid plans for advanced features.

How do I get started with PageIndex?

Visit https://pageindex.ai/ to sign up and explore PageIndex.

Explore more AI Ai Services tools

Browse all Ai Services tools →

Compare PageIndex: vs oMLX · vs Olostep · vs FreeLLMAPI · vs Speko