Donut: OCR-Free Document Understanding Transformer
github.com
github.com
I think the traditional approach to scanning and classifying without AI/ML is the way to go, for the next 5 years at very least.
For my use cases, this already beats all "traditional approaches" for at least a few month now. That's just inferring from when I first stumbled across it. No clue for how long it's been a thing.
Which it used OCR to produce digital text.
So one source of training data at least.
Leaking sensitive data of enterprise customers as training material for public recaptchas falls in that category.
Google Vision: 95.62% HW - 99.4% Typed
Amazon Texttract: 95.63% HW - 99.3% Typed
Azure: 95.9% HW - 98.1% Typed
Then if curious, TrOCR was the best FOSS solution at 79.5% HW and 97.4% Typed. (However it took roughly 200x longer than Tesseract which was 43% HW and 97.0% Typed)
I also tested paddleocr and keras ocr to round them all out.
At some point I really need to finish my project enough to write up some blog articles and post a bunch of code repos for others to use.
It uses Tesseract under the hood. Results tend to just be OK in my experience.
I'll note that when I put the tesseract output into chatgpt and prompted it saying it was ocr'd text and asking to clean it up, it worked very well.
My first time processing it, I used `ocrmypdf --redo-ocr` because it looked like there was some existing OCR. After processing, the OCR was crap because ocrmypdf didn't realize it was OCR but thought it was real text in the document that should be kept. This was fixable using `ocrmypdf --force-ocr`.
Before realizing this, I discovered that Tesseract 4 & 5 use a neural network-based recognition. I then came across this step-by-step guide on fine-tuning Tesseract for a specific document set: https://www.statworx.com/en/content-hub/blog/fine-tuning-tes...
I didn't end up following the fine-tuning process because at this point `ocrmypdf --force-ocr` worked excellently, but I thought the draw_box_file_data.py script from their example was particularly useful: https://gist.github.com/flaviut/d901be509425098645e4ae527a9e...
I did a presentation on the topic recently: https://clis-everywhere.k8s.best/16
I'll soon make the stack open source, but it shouldn't be hard to recreate given the inputs I've already provided.
But this isn’t much help if you must classify images.
This seems like exactly the kind of problem that will see rapid improvements as people point more LLMs at multimodal input.
Right now making predictions for ML capabilities on a five year timeframe seems foolhardy.
Smaller graph models could get away with using less data. The problem that the "traditional" approach had is the the quality of the OCR was the bottleneck for overall model performance. It amazes me how this problem shifted from a node classification problem to a image to text problem.
Training on CPU was possible with GCN but not with Donut.
Author: phew! I'm glad there's an 'n' in there somewhere
https://chat.openai.com/share/25124525-0bad-4c13-ae5a-ae4bea...
3.5 doesn't really get it, 4 does. There are some that aren't great but the context is pretty small and it can be a decent launching point.
https://cloud.typingmind.com/share/b49794e0-4b2b-4213-ae2e-8...
I'll grant one thing: the paper must go well with cheese.
It reminds me of the microwave cookery books that came out after consumer microwaves became available: there are things a microwave is good at, but those books used it for everything, just like we're using GPT-4 today. We'll calm down eventually.
...but he's using a language model for a language task.
That's exactly how many Windows office users do it, they paste a list into Excel and use it to remove duplicates. There are alternatives even on Windows, but It's must easier for them to use a single general-purpose graphical tool (and let's not get started with the abominations VLOOKUP is used for).
I used to look down on that but then I realized that using a graphical program for list manipulation is kinda cool and that this program is rather capable, and could create combinations that are rather difficult to do with the more specialized tools. I still use these specialized tools (I'm used to them, and I can do some stuff they can't easily do in Excel).
Yes, I know sort | uniq. I even have a couple Linux shells open on my Windows work system. But I can't for the life of me remember the magic flags, so I'll either paste the list to Emacs and M-x sort-lines + M-x delete-duplicate-lines, or paste it to Excel and do it there, or do something even more cheesy - whatever is least likely to break my flow.
There are tools more or less optimized for any specific job, but the best tool for the job is the one you have handy, and are experienced in using.
I too increasingly often find myself using GPT-4 for random, ad-hoc tasks. They may or may not be better tools out there. I may even have some installed. But none of them beat being able to just describe what you want, paste some data, and get the results out few seconds later.
Some of the answers are straight up usable, others if you prefer you can go from there because this a creative language task.
And there isn't really a specific tool for this, is there? It's nothing like your comparison to a very well specified problem. "Identify what this thing does, and come up with a title that also contains a word, and the word is related to the topic" is not the same as sort|uniq Vs a spreadsheet.
(1) The problem is not clearly defined. Does the word need to be thematically related to the topic? (As far as I can tell, "Donut" isn't thematically related to document understanding.) Maybe you could say it's a nice, optional bonus if it's related.
(2) The best solution would be good at two things: (A) satisfying constraints and (B) creativity. ChatGPT is unlikely to be good at A, and a non-AI algorithm that just finds valid words can't do B.
Regarding #1, if people don't all have the same idea of the problem, they're not going to agree on the solution.
Regarding #2, maybe a combined solution would be best. Generate all allowable words, then feed them to ChatGPT and have it say which ones are thematically good.
These are hard to encode.
Instead I just asked "Make them more fun, and related to literary characters", then Muppets and awkward ones based on Harry potter which it described as "certainly a unique request". It's faster than getting a word list related to that. And they are frankly great - better than I'd come up with given much longer.
https://chat.openai.com/share/d9a31442-96e5-4a18-92d9-d217c7...
This problem is great for llms. It's language, works well with a back and forth discussing good and bad options, has no well defined output requirements but is easy to explain to a person, has a human in the loop and has almost zero cost if it's wrong.
> QUIXOTE: Quality Unstructured Information Extraction and Organization Through End-to-end transformer
If you want to do it, using “plant based” is probably better than “vegan”, and it’s always good to make sure your users are aware that the mark can be wrong and they should double-check with the waiter.
As for your question - I didn’t play with Donut, but ocr+gpt or multimodal gpt4 once released should handle this smoothly.
You could combine ingredient search, looking for symbols the actually designate vegan as some places do, along with long/lat data to determine what the restaurant actually is and then check it with a database you maintain.
So I could scan a menu, and then ask the owner or server about certain dishes, and then crowdsource an updated
It would be great if there was a standard API for all restaurants that included all menu items, prices, ingredients, preparation and sourcing information. I could be maintained like a wiki I suppose, and restaurants could be incentivized by including their restaurants.
Trying to solve the problem of scanning through menus for multiple restaurants to find something a vegan or vegetarian can eat, and instead just showing all the individual menu options in the area as a list.
The only problem is that the data is closely walled by Google and you can only access it through their api.
What I want to create is a tiny search engine that collects all menus (somehow collected from images) and let users find and filter what they like, and even get recommendations nearby!
Perhaps if iOS would allow an AI to access the screen pixels directly ...
The training process is not hard, but the data gathering / cleaning / labelling can be a little long.
Some users have expected it to "unwarp" bad scans, it also doesn't do that unfortunately. But that's a much harder problem to solve...
I'm curious as to why this appears in your ToS, as this is quite a deterrent for many.
Feels like someone trying to throw a stake in the ground rather than releasing a quality product, honestly.
It's surprisingly accurate although you have to write your own program to segment the image into lines. I think with some fine tuning I could have the machine read my notebook with minimal corrections.