HNHacker News
TopNewBestAskShowJobs

mingtianzhang

193 karma · joined February 16, 2023

personal website: mingtian.ai
submissionscomments
mingtianzhang··on Show HN: Open KB: Open LLM Knowledge Base
Hi, thanks for the feedback. Our biggest differentiation is that OpenKB can handle long PDFs and images—something that isn’t trivial or easily “vibe-coded.” We will try to find some ways to compare with other projects! Thanks for the valuable feedback!
mingtianzhang··on [dead]
VLM can already process both the document images and the query to produce an answer directly. Do we still need the intermediate OCR step?
mingtianzhang··on Do we still need OCR? An implementation of a pure vision-based agent
We discuss the limitations of the classic OCR pipeline and provide a pure vision-based RAG system for document analysis (https://github.com/VectifyAI/PageIndex/blob/main/cookbook/vi...)

Any feedback is welcome!

mingtianzhang··on Should LLMs just treat text content as an image?
We actually don't need OCR: https://pageindex.ai/blog/do-we-need-ocr
mingtianzhang··on Do We Still Need OCR?
This blog examines the inherent limitations of the current OCR pipeline in the context of document question-answering systems from an information-theoretic perspective and discusses why a direct, vision-based approach can be more effective. It also provides a practical implementation of a vision-based question-answering system for long documents.
mingtianzhang··on Reasoning-based RAG for long document question answering
PageIndex Chat is the world's first human-like long-document AI analyst. You can upload entire books, research papers, or hundred-page reports and chat with them without context limits, all in the browser. Unlike traditional RAG or "chat-with-your-doc" tools that rely on vector similarity search, PageIndex builds a hierarchical tree index of your document (like a table of contents), and then reasons over this index to retrieve and interpret relevant sections. It doesn’t search by keywords or embeddings — it reads, understands, and reasons through the document like a human expert.

What makes it different:

- Reasoning-based retrieval: Understands structure, logic, and meaning, not just semantic similarity.

- Page-level references: Every answer includes precise citations for easy verification.

- Cross-section reasoning: Connects information across sections and appendices to find true answers.

- Human-in-the-loop: You can guide, refine, and verify its reasoning.

- Multi-document comparison: Analyze and contrast multiple reports at once.

mingtianzhang··on PageIndex Chat – Human-Like Long Document AI Analyst
PageIndex Chat is the world's first human-like long-document AI analyst. You can upload entire books, research papers, or hundred-page reports and chat with them without context limits, all in the browser.

Unlike traditional RAG or "chat-with-your-doc" tools that rely on vector similarity search, PageIndex builds a hierarchical tree index of your document (like a table of contents), and then reasons over this index to retrieve and interpret relevant sections. It doesn’t search by keywords or embeddings — it reads, understands, and reasons through the document like a human expert.

What makes it different:

- Reasoning-based retrieval – Understands structure, logic, and meaning, not just semantic similarity. - Page-level references – Every answer includes precise citations for easy verification. - Cross-section reasoning – Connects information across sections and appendices to find true answers. - Human-in-the-loop – You can guide, refine, and verify its reasoning. - Multi-document comparison – Analyze and contrast multiple reports at once.

mingtianzhang··on PageIndex Chat – Human-Like Long Document AI Analyst
PageIndex Chat is the world's first human-like long-document AI analyst. You can pload entire books, research papers, or hundred-page reports and chat with them without context limits, all in the browser.

Unlike traditional RAG or "chat-with-your-doc" tools that rely on vector similarity search, PageIndex builds a hierarchical tree index of your document (like a table of contents), and then reasons over this index to retrieve and interpret relevant sections. It doesn’t search by keywords or embeddings — it reads, understands, and reasons through the document like a human expert.

What makes it different:

- Reasoning-based retrieval – Understands structure, logic, and meaning, not just semantic similarity. - Page-level references – Every answer includes precise citations for easy verification. - Cross-section reasoning – Connects information across sections and appendices to find true answers. - Human-in-the-loop – You can guide, refine, and verify its reasoning. - Multi-document comparison – Analyze and contrast multiple reports at once.

mingtianzhang··on DeepMind's paper reveals Google's new direction on RAG: In-Context Retreival
Instead of relying on vector databases, DeepMind proposes:

1. The LLM itself selects the most relevant documents — no vector database needed.

2. The selected documents are then placed directly into the context for generation.

This kind of in-context retrieval approach greatly improves retrieval accuracy compared to traditional vector-based retrieval methods.

mingtianzhang··on Show HN: A Vectorless LLM-Native Document Index Method
Hi, thanks for your inspiring questions.

1. What happens when the TOC is too long? -- This is why we choose the tree structure. If the ToC is too long, it will do a hierarchy search, which means search over the father level nodes first and then select one node, and then search its child nodes.

2. How does the index handle near misses, and how do you disambiguate between close titles? For each node, we generate a description or summary to give more information rather than just titles.

3. For documents that are not in a hierarchy, it will just become a list structure, which you can still look through.

We also write down how it can combine with a reasoning process and give some comparisons to Vector DB, see https://vectifyai.notion.site/PageIndex-for-Reasoning-Based-....

We found our MCP service works well in general financial/legal/textbook/research paper cases, see https://pageindex.ai/mcp for some examples.

We do agree in some cases, like recommendation systems, you need semantic similarity and Vector DB, so I wouldn't recommend this approach. Keen to learn more cases that we haven't thought through!

mingtianzhang··on Which table format do LLMs understand best?
The current OCR approach typically relies on a Vision-Language Model (VLM) to convert a table into a JSON structure. However, a table inherently has a 2D spatial structure, while Large Language Models (LLMs) are optimized for processing 1D sequential text. This creates a fundamental mismatch between the data representation and the model’s input format.

Most existing pipelines address this by preprocessing the table into a linearized 1D string before passing it to the LLM — a question-agnostic step that may lose structural information.

Instead, one could retain the original table form and, when a question is asked, feed both the question and the original table (as an image) directly into the VLM. This approach allows the model to reason over the data in its native 2D domain, providing a more natural and potentially more accurate solution.

mingtianzhang··on Show HN: Long PDF Reader MCP
Thanks, any feedback is welcome!
mingtianzhang··on Managing context on the Claude Developer Platform
Thanks for the reminder, I have edited the comment.
mingtianzhang··on Managing context on the Claude Developer Platform
Edited version:

We try to solve a similar problem to put long documents in context. We built an MCP for Claude to allow you to put long PDFs in your context window that go beyond the context limits: https://pageindex.ai/mcp.

mingtianzhang··on Show HN: Long PDF Reader MCP
Thanks for the great question! We actually use a reasoning-based, vectorless approach. In short, it follows this process:

  1. Generate a table of contents (ToC) for the document.

  2. Read the ToC to select a relevant section.

  3. Extract relevant information from the selected section.

  4. If enough information has been gathered, provide the answer; otherwise, return to step 2.
We believe this approach closely mimics how a human would navigate and read long PDFs.
mingtianzhang··on Show HN: Long PDF Reader MCP
thanks!
mingtianzhang··on Event Sourcing, CQRS, and Microservices: A Real FinTech Example from My Career
Great post, I am wondering if this system includes financial report analysis.
mingtianzhang··on Handle in-document reference in RAG
Many RAG systems handle in-document references (like “see appendix for details”) by building graphs or other preprocessing structures. The idea is to make sure cross-references are resolved before retrieval.

But with reasoning-based RAG, you don’t need that extra layer. The LLM itself can read the document, notice the reference, and then “jump” to the appendix (or wherever the reference points) to extract the answer. In other words, instead of pre-building structure, the model reasons its way through the content.

An example of reasoning-based RAG with PageIndex MCP is attached. In this example, the query asks for the total value. The main text only provides the increased value and refers to the appendix table for the total value. The LLM then looks up the appendix to find the total value and explains its reasoning process.

This raises an interesting question: how much preprocessing do we actually need for reasoning-augmented RAG, and when is it better to just let the model figure it out?

mingtianzhang··on Fine-Tune Black Box Embedding Models
This paper introduces a method that allows you to fine-tune black box embedding models (e.g. those vectors obtained with ChatGPT API). It shows there is around 10% improvement in various domains. Any feedbacks are welcome.
mingtianzhang··on Show HN: I built an AI in 3 days at 16 y/o that lets you chat with your PDFs
Check out this MCP: https://pageindex.ai/mcp, which allows you to chat with any long PDFs (hundreds of pages) beyond the context limit of Claude or ChatGPT.
mingtianzhang··on [dead]
The word index originally came from how humans retrieve information: book indexes and tables of contents that guide us to the right place.

Computers later borrowed the term for data structures such as B-trees, hash tables, and more recently, vector indexes. They're highly efficient for machines, but also abstract and unnatural: not something a human, or an LLM, can directly use as a reasoning aid. This creates a gap between how indexes work for computers and how they should work for models that reason like humans.

PageIndex is a new step that looks back to move forward. It revives the original, human-oriented idea of an index and adapts it for LLMs. Now the index itself (PageIndex) lives inside the LLM's context window: the model sees a hierarchical table-of-contents tree and reasons its way down to the right span, much like a person would retrieve information using a book's index.

PageIndex MCP shows how this works in practice: it runs as a MCP server, exposing a document's structure directly to LLMs. This means platforms like Claude, Cursor, or any MCP-enabled agent can navigate the index themselves and reason their way through documents, not with vectors or chunking, but in a human-like, reasoning-based way.

mingtianzhang··on Show HN: Chat with Long PDFs on Cursor or Claude Desktop
Hi, we generate a table of contents (ToC) for each document, so LLM can use the ToC to navigate the long documents, which can bypass the context limit. You can checkout this notebook for a quick toturial about our method: https://docs.pageindex.ai/cookbook/vectorless-rag-pageindex
mingtianzhang··on Defeating Nondeterminism in LLM Inference
I agree that we need stochasticity in a probabilistic system, but I also think it would be good to control it. For example, we need the stochasticity introduced at high temperatures since it is inherent to the model, but we don’t need stochasticity in matrix computations, as it is not required for modeling.
mingtianzhang··on Voyager – An interactive video generation model with realtime 3D reconstruction
What's your opinion on modeling the world? Some people think the world is 3D, so we need to model the 3D world. Some people think that since human perception is 2D, we can just model the 2D view rather than the underlying 3D world, since we don't have enough 3D data to capture the world but we have many 2D views.

Fixed question: Thanks a lot for the feedback that human perception is not 2D. Let me rephrase the question: since all the visual data we see on computers can be represented as 2D images (indexed by time, angle, etc.), and we have many such 2D datasets, do we still need to explicitly model the underlying 3D world?

mingtianzhang··on 'World Models,' an old idea in AI, mount a comeback
I used to work on a idea that instead of modelling the whole world, you can build your own Solipsistic model: https://openreview.net/pdf?id=fPaGSuQRP1O
mingtianzhang··on SynthID – A tool to watermark and identify content generated through AI
Do you think if such a tool exists, will it benefit the community or not?
mingtianzhang··on The Theoretical Limitations of Embedding-Based Retrieval
"why would we want an LLM do something as inefficiently as a human?" -- That is a good point. Maybe we should rename artificial intelligence (AI) to super-artificial intelligence (SAI).
mingtianzhang··on SynthID – A tool to watermark and identify content generated through AI
Yeah exactly, you can always do that by using another model that doesn't have the watermark.
mingtianzhang··on The Theoretical Limitations of Embedding-Based Retrieval
Just similar in theorem style, I try to emphasise that no lossy representation is universally (i.e. for all downstream tasks) better than another.
mingtianzhang··on SynthID – A tool to watermark and identify content generated through AI
Based on my research experience and judgment, I have published several top-conference papers in both the detection and diffusion domain, but I haven’t explored the engineering/product side. I believe that if such a system hasn’t been invented yet, it wouldn’t be difficult to create one to remove that watermark using an open-source image/video model and maintain the high quality. Would you be interested in having a further discussion on this?
Page 1 of 3Next →