Full threadmingtianzhang·VLM can already process both the document images and the query to produce an answer directly. Do we still need the intermediate OCR step?View on HN