Benchmarking vision-language models on OCR in dynamic video environments
arxiv.org
arxiv.org
In Figure 1 the authors complain that Gemini “misreads 'ss ety!' as 'ness ety!'”, but even a casual look at the image reveals that Gemini's reading is correct.
In Figure 11, they state that Claude is “altering the natural sequence of ideas in the ground truth”, except that the sequence in the ground truth makes no sense, while Claude's order does (only the initial “the” is misplaced).
TBH, I'm not sure it's a good test. I can somewhat see the argument against "BASELINE" for ground truth - the underlying text might have been BASE(IAKS), for all we know. But, IMO the ground truth should have been "Direction & ess" at the very least. And, more significantly than that - it's a fake scenario, that we don't care for in practice. Why use that? Use invoices with IDs that sound like words but are not. Use license plates and stuff like that. Heck, use large prints of random characters, mixed with handwritten gibberish.
For at least some of images that they used, the expectation from a good text reader is actually to understand context and not blindly OCR. Take "Trader Joe's": we *know* that's an 's', but only from outside context; from OCR, it might've been an 8, there's really no way to tell. Why accept the "s" in ground truth, but reject the full world "Coconut" (which is obviously what is written on the can, even if partially obscured)? Furthermore, a human would know what kind of products are sold by Trader Joe's, and coupling that with the top of the letters "M I L" that are visible, would deduce that's Coconut Milk. So really, Claude nailed that one.
What reviewers?
EasyOCR is LSTM-CTC from 2007, RapidOCR is a ConvNet approach from 2021, both focused on speed. Both will vastly outperform almost any transformer model, and certainly a big one, on speed and memory usage, but they aren't state of the art on accuracy. This is well known, for a decade at this point. 2 decades for LSTM-CTC.
Plus, I must say the GPT-4o results look a lot saner. "COCONUT" (GPT-4o) vs "CONU CNBC" (Gemini) vs Ground Truth "C CONU CNBC". And, obviously the ground truth should be "COCONUT MILK" (the word milk is almost entirely out of the picture, but is still the right answer that a human would give). The "C CONU" comes from the first O of COCONUT being somewhat obscured by a drawing of ... I don't know what the hell that is. It's still very obvious it's meant to be "COCONUT MILK", so the GPT-4o answer is still not quite perfect, but heaps better than all the others.
Now this looks very much like it might be temperature related, and I can find nothing in the paper about changing the temperature, which is imho a very big gap (temperature gives transformer models more freedom to choose more creative answers. The better performance of GPT-4o might well be the result of such a more creative choice, and might also explain why Gemini is trying so hard to stay so very close to the ground truth. It's still quite the accomplishment to succeed, but GPT-4o is still better)
It'll also depend if you care about tabular data, whether a 'minor' numerical error (like 0 & 8 mismatched sometimes) is significantly worse than a 'typo' as it were in recognising a word, etc.
Overall | Handwritten | Typed
Google Vision: 98.80% | 93.29% | 99.37%
Amazon Texttract: 98.80% | 95.37% | 99.15%
surya: 97.41% | 87.16% | 98.48%
azure: 96.09% | 92.83% | 96.46%
trocr: 95.92% | 79.04% | 97.65%
paddleocr: 92.96% | 52.16% | 97.23%
tesseract: 92.38% | 42.56% | 97.59%
nougat: 92.37% | 89.25% | 92.77%
easy_ocr: 89.91% | 35.13% | 95.62%
keras_ocr: 89.7% | 41.34% | 94.71%
Handwritten is a weighted average of Handwritten and typed, I also did Jaccard and Levenshtein distance, but the results were similar enough that just leaving them out for sake of space.Overall, of you want the best, if you're an enterprise, just use whatever AWS/GCP/Azure you're on, if you're an individual, pick between those. While some of the Open Source solutions do quite well, surya took 188 seconds to process 88 pages on my RTX 3080, while the cloud ones were a few seconds to upload the docs and download them all. But if you do want open source, seriously consider surya, tesseract, and nougat depending on your needs. Surya is the best overall, while nougat was pretty good at handwriting. Tesseract is just blazingly fast, from 121-200 seconds depending on using the tessdata-fast or best, but that's CPU based and it's trivially parallelizeable, and on my 5950X using all the cores, took only 10 seconds to run through all 88 pages.
But really, you need to generate some of your own sample test data/examples and run them through the models to see what's best. Given frankly how little this paper tested, I really should redo my study, add VLMs, and write a small blog/paper, been meaning to for years now.
Maybe? Seems application-dependent to me.
If you're OCRing checks or invoices or car license plates or tables in PDF documents, you might prefer a model that's more conservative when it comes to filling in the blanks!
And even when recognising packaged coconut products, you've also got your organic coconut oil, organic coconut milk with reduced fat, organic coconut cream, organic coconut flakes, organic coconut dessicated chips, organic coconut and strawberry bites, organic coconut milk powder, organic coconut milk block, organic coconut milk 9% fat, organic coconut yoghurt, organic coconut milk long life barista-style drink, organic coconut kefir, organic coconut banana and pear baby food pouches, organic coconut banana and pineapple smoothie, organic coconut scented body wash and so on.
You also didn’t really counter the paper. Sure, the OCR models are old, but what should they have tested instead? Are there better open-source OCR models available that would have made for a fairer comparison?
It's clearly the stem from the bell pepper in front of the can. You're complaining that the software is lesser than a human, yet it appears your human needs better training in understanding context too.
Edit to insert: WHAT DRAWING? There's a can of coconut milk that is turned so the word coconut is not fully visible. In front of that can is a real red bell pepper with a green stem still attached that is partially obstructed by the bowls in the foreground. What you're attempting to claim as a drawing is just a real life object in the table top setup. Since this is a CNBC branding image, I'm assuming this is a still frame from a video clip. Based on being a video type person, this view probably changes based on time with different things being obstructed/revealed by the camera's movement.
Your RLHF could really use some improvement. To be this argumentative when you're clearly wrong is quite amusing, but not in an entertaining way. It just reinforces my sentiments towards the joke the industry has become
[1]: https://github.com/opendatalab/MinerU [2]: https://github.com/opendatalab/OHR-Bench [3]: https://github.com/reductoai/rd-tablebench
> Three state of the art VLMs - Claude-3, Gemini-1.5, and GPT-4o
Literally none of those are state of the art. Academia is completely unprepared to deal with the speed Ai develops. This is extremely common in research papers.
That's literally in the abstract. If I can see a completely wrong sentence 5 seconds into reading the paper, why should I read the rest?
"Models from leading AI labs" or similar. Leaving it like now signals either sloppiness or dishonesty
Honestly I thought Claude-3 and GPT-4o were some of the newest major models with vision support, and that models like o1 and deepseek were more reasoning-oriented than OCR-oriented.
I'm not that familiar with Claude for vision. I don't think Anthropic focusses on that. But the 3.5 family of models is way better. If 3.5 Sonnet supports vision that's what I'd use
It was literally launched February 5th, ~10 days ago. I'm no researcher, and I know "academia moves slow" is of course true too, but I don't think we can expect research papers to include things that were launched probably after they finished the reviews of said paper.
Maybe papers aren't the right approach here at all, but I don't feel like it's a fair complaint they don't include models released less than 2 weeks ago.
Also, this is arxiv. The website that's explicitly about posting research pre peer-review.
So for how long? How long did the papers you've written in the past take to write? AFAIK, it takes some time.
And peer-review is not the only review a paper goes through, and was not the reviews I was referring to.
I'm not even disagreeing that it takes time to write papers, and it's "common" for this to happen. But it's just more evidence for what I said in my original comment:
> Academia is completely unprepared to deal with the speed AI develops
The real challenge is building reliable and accurate ETL pipelines (document ingestion from web, OCR, classification, validation, etc.) that work at scale in production.
The best products will be defined by everything "non-AI", like UX, performance, and human-in-the loop feedback loop for non-techies.
Avoiding over-reliance on specific models also helps. With good internal eval data and benchmarks, you can easily switch or fine-tune models.
By building a good UX and integrating it with other processes that require traditional collaboration, you increase the chances that replicating your secret sauce is either infeasible or too difficult for newcomers to bother.
The initial results were quite promising, as GPT-4o could reliably identify the correct place in the form for the information, and moderately reliably extract the values, even if the image was blurry or the text was sloppily written. Excited to see how Gemini 2.0 would do on this task!
[0] https://arxiv.org/abs/2412.15260
[1] https://github.com/hwestermann/AI4A2J_analyzing_images_of_le... (code and data)
It feels to me that if you need to provide schema and preprocess the data and this and that at the end all AI provide is a way to do some SQL in natural language, meaning yes it's better but it doesn't remove the actual pain point if you're a tech user.
Then again maybe I'm wrong, didn't find the right tool or didn't understand it.
Is what I'm looking for something that actually exists (and works, not just on simple cases)?
We tried several dedicated services for extracting structured data and factoids like that from documents: First Google Document AI, then a dedicated provider focusing solely on our niche. Back then, that gave the best results.
There wasn't enough budget to go deeper into this and we just reverted to doing it manually. But I think a really cool way to do this would be to make a user friendly UI where they can see suggestions and the text snippets they were extracted from as they skim through the document, with a simple way to modify and accept these. I think that'd work to scale the process quite a bit. Focusing the attention of the human at the relevant parts of the document basically.
Haven't worked on this space since then, but I'm pretty bearish on fully automated fact extraction. Getting stuff in contracts and invoices wrong is typically not acceptable. I think a solid human in the loop approach is probably still the way to go.
Between this and their 1M token context it is getting hard to ignore Google's models.
"Perform OCR on this image. Return only the text found in the image as a single continuous string without any newlines, additional text, or commentary. Separate words with single spaces. For any truncated, partially visible, or occluded text, include only the visible portions without attempting to complete or guess the full text. If no text is present, return empty double quotes."
Found in: https://github.com/video-db/ocr-benchmark/blob/main/prompts....
Yet another paper where the authors don't address what tokens are. It's like publishing Rolling pin fails at math or Calculator fails to turn dough ball into round pizza.
While I can understand where they're coming from in a desire to avoid hallucination when doing some letter for letter transcription from an image, certainly most times you reach for OCR you want the original copy, despite damage to its representation (paper tears, coffee stains, hands in front of it). Turns out token conjunction probability conjectures come in handy here!
Whether the image of an object, or the object, is "Ground Truth" is an exercise left to the user's goal. Almost all use cases would want what was originally written on the object, not its present occlulded [sic] representation.
Obviously these models would have lower accuracy, but running at all would be nice.
https://news.ycombinator.com/item?id=43048326
Throw those into Google search along with term iOS or Android.
Excellent Freudian slip (proverb allusion suggesting Google has a blind spot, while discussing OCR).