That said, there are some large language models you can run locally which accept image input. Phi-3-Vision [3], LLaVA [4], MiniCPM-V [5], etc.
[1] - https://github.com/Dicklesworthstone/llm_aided_ocr/blob/main...
[2] - https://github.com/ggerganov/llama.cpp?tab=readme-ov-file#de...
[3] - https://huggingface.co/microsoft/Phi-3-vision-128k-instruct
Although LLaVA specifically it might not be great for OCR; IIRC it scales all input images to 336 x 336 - meaning it'll only spot details that are visible at that scale.
You can also search on HuggingFace for the tag "image-text-to-text" https://huggingface.co/models?pipeline_tag=image-text-to-tex... and find a variety of other models.
The latest architecture is supposed to improve this but there are better architectures if all you want is OCR.
https://huggingface.co/xtuner/llava-llama-3-8b-v1_1-gguf
But I see that this new one just came out using Llama 3.1 8B:
https://huggingface.co/aimagelab/LLaVA_MORE-llama_3_1-8B-fin...