777 karma · joined January 7, 2011
Agreed the dimensions/features are key. These white papers are an insult to science...
I see OCR much like phonemes in speech, once you have end to end systems, they become latent constructs from the past.
And that is actually good, more code going into models instead.
For a VLLM, my understanding is that OCR corresponds to a sub-field of questions, of the type 'read exactly what's written in this document'.
On a side note, there's decent research on how well bilingual humans do actually think in both language, and are actually better at decisive thinking outside of their mother tongue.
Raw code is here: https://gitHub.com/beniz/llmbox
All runs locally, the finetune is a plaigemma-3b (from mix-448). Acc/F1/prec/recall are all within 99.99%, including on llm-generated spam.
It's a reduction of what AI is as a computer science field and even of what the subfield of generative AI is.
On a positive note, generative AI is a malleable statiscally-geounded technology with a large applicative scope. At the moment the generalistic commercial and open models are "consumed" by users, developers etc. But there's a trive of forthcoming, personalized use cases and ideas to come.
It's just we are still more in a contemplating phase than a true building phase. As a machine learnist myself, I recently replaced my spam filter with a custom fineruned multimodal LLM that reads my emails a pure images. And this is the early early beginning, imagination and local personalization will emerge.
So I'd say, being tired of it now is missing much later. Keep the good spirit on and think outside the box, relax too :)