Llama 3-V: Matching GPT4-V with a 100x smaller model and 500 dollars
aksh-garg.medium.com
aksh-garg.medium.com
You need to compare the evals to strong open VLMs including this and CogVLM
- This is not "first-ever multimodal model built on top of Llama3", there's already a Llava on Llama3-8b https://huggingface.co/lmms-lab
(InternVL appears to hallucinate more.)
People upvoting the post??
Not really sure? But PT Barnum said there's always a lot of them out there.
Pretty sure they mean fine tuning though?
But even that is total tripe.
These guys are snake oil salesmen. (Or Sylvester McMonkey McBean is behind it.)
> We add a vision encoder to Llama3 8B
https://news.ycombinator.com/item?id=39136472
Many companies also optimize for tools, like Python, that have boost productivity more than price/performance ratio. OpenAI had billions of other people's money. They might just keep using tools which worked before.
Lastly, there are tons of papers published on techniques that claim to reduce cost. Most of them aren't good. Their benchmarks aren't good. Even reviewing most of them is more time than a lot of AI researchers have. Those that make it to established communities usually have gotchas that come with the benefits. So, they could also simply miss a needle in a large haystack.
I think you're right that they'd be using whatever really worked with no loss in model performance. It's just that they might not for a number of reasons. The rational choice is for others to keep experimenting with those things in case they get a competitive advantage.
The first time a model that actually matched GPT 4 launched (i.e. Command-R+) there was no mention of it at all. If your results speak for themselves, there's no need to shout.
Edit: Thank you!
However Tesseract is quite behind still as you note, even with v5.
I would say PaddleOCR is good in general for tables - it's much better (in terms of recall rate) at recognising numerical digits / symbols than Tesseract although I notice it often misrecognises "l" in "Lullaby/ml/million" etc as "1" sometimes.
The cloud providers have better table extraction iff you can guarantee the same format each time for the document.
Here's just one example: https://www.totalflood.com/samples/residential.pdf (I struggle getting accurate data out of the Sales Comp section - basically all approaches mix up the properties.
I would suggest your best bet is waiting 2 years for the next version of LLAVA to come out which may have capabilities to interpret very accurately on device. The progress with LLAVA has been fast recently but for now it's still a bit too inaccurate.
Does this require a slight offset and/or rotation to the image, or just literal rerun, with seed seed/whatever giving a different result?
As in: you tell it that these and these parts should be masked such and such, and then it does that?
And these haven't even been trained to defeat captchas/logic problem captchas yet, if it was fine tuned on the general pattern of them I imagine any form of captcha is bust.
I don't see how an older smartphone could meaningfully outcompute a spamming infra.
Of course eventually this will be defeated too, but for now it seems to work pretty well.
[0] https://github.com/OpenBMB/MiniCPM-V
[1] https://huggingface.co/spaces/openbmb/MiniCPM-Llama3-V-2_5
You only need an email address.
https://build.nvidia.com/microsoft/phi-3-vision-128k-instruc...
Traditional OCR do not handle multiple invoice formats or handwritten ones.
I would like to train one locally with as many invoices it wants
CogAgent is also CogVLM modified to handle documents and larger images. CogVLM is better for VQA.
This absolutely is not the first Llama3 vision model. They even quote it's performance compared to Llava. Hard to take anything they say seriously with such obviously false claims
Although this is true, there have been earlier Llama3 based vision releases, none of the latest Llava releases are Llama3 based.
Llama 3 outputs text and can only see text, this is a vision model.
>that would make it Llama-2-based.
It's based on Llama 3, Llama 2 has nothing to do with it. They took Llama 3 Instruct and CLIP-ViT-Large-patch14-336, train the projection layer first and then later finetuned the Llama 3 checkpoint and train a LoRA for the ViT.
It is not by the original group who have published a series of models under the Llava name.