Do we still need OCR when we can build a pure vision-based AI agent
pageindex.ai
pageindex.ai
However, this paradigm shift raises an important question:
If a VLM can already process both the document images and the query to produce an answer directly, do we still need the intermediate OCR step?