Good points.
>> no matter how I scan receipts or documents, the OCR output seems far worse than human
> what kind of errors are you getting?
Just noticeably worse character recognition, particularly where the document is faded, water-stained, or the paper document (not the scan) was low resolution to begin with, as compared to 'normal' human recognition.
I was largely using Tesseract in conjunction with the self-hosted Paperless-NGX, and I wanted to stay free/local without yet investing in AI-focused hardware. But you're right that AI will continue to advance, including in smaller local models, and it looks like Paperless 3.0 released recently (after my testing earlier in 2026), including with AI functionality.
> or having the documents smooshed together because the OCR can't parse the layout
I haven't even been worrying about that yet - I'm just at the point of trying to get the OCR characters right. :-)