TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
github.com
github.com
It stands out by requiring merely a 24G GPU for training and an 8G GPU or CPU for inference. Built upon Phi-2, TinyGPT-V couples an effective language backbone with pre-trained vision modules from BLIP-2 or CLIP. TinyGPT-V's 2.8B parameters can undergo a unique quantisation process, suitable for local deployment and inference tasks on 8G various devices.
On my M1 Macbook Air, current LLMs can run surprisingly quickly locally already.
Agreed. The huge leap-forward in capability already boggles my mind, the fact that I can run it _on my desktop cpu_ and have an interactive, natural-language oracle of the internet in a mere ~11GB file.
Everyone seems to be chasing the network-accessible API approach because lock-in is easy and if you're at the bleeding-edge (training) you've got the compute for running it anyway.
But now with accessible, local models my bet is on bored young hackers coming up with the best use cases and killer apps, not Microsoft.
This feels like a lazy critique. Sure companies may like the lock-in, but people are chasing the network-accessible API approach because it is way more powerful than the stuff you can run locally. Admittedly it’d be great if local models could get up to ChatGPT’s level, but if we’re being honest with ourselves they’re not there yet and are miles away from GPT-4, so it’s totally understandable that people want what OpenAI is selling.
That's only one part of the equation because these young hackers need companies with huge capitalization to deliver the models to them for free. What I hope actually happens is that model building can become modular and therefore could be crowdsourced. I'm not even sure this is even theoretically possible but it would help a lot.
[1]: https://huggingface.co/microsoft/phi-2/blob/main/LICENSE
Am I missing something here? Did the authors forget about for loops? What happens if you only do it 16 times?
I should really have got over it after 40 years.
Explain that!
you can’t see what’s not there.
you aren’t a language model that ate after midnight.
I am with you. The movie has no hints or clues when it's okay to feed again. Dawn ? Sunrise ?
wait, unless you were joking in the first place.
Most professionals in the field that I'm near would not touch that model with a 10 foot pole. We desperately need better validation/data contamination detection methods.
Unless there was something else for phi-2 as well?
the top of the HuggingFace Open LLM leaderboard was pretty meaningless for a bit because many of the top scores were achieved by training on the evaluation data
https://huggingface.co/spaces/Qwen/Qwen-VL-Plus
You can also ask it to give you bounding boxes of objects.
The papers only have one common benchmark (GQA, MobileVLM scores better) so hard to say how they compare otherwise.
-
EDIT - GPT helped me understand better:
--
>>> "This model is special because it can do similar tasks as the big models but requires much less computational power1. It’s like having a small but powerful engine that can do the work of a big one. This makes it more accessible for more people to use it"
---
>>> "TinyGPT-V is built on another model called Phi-2 and uses pre-trained vision modules from BLIP-2 or CLIP1. It has 2.8 billion parameters (these are like the model’s brain cells) and can be further compressed to fit on devices with 8GB memory1. This means you could potentially run this model on your personal computer or even some high-end smartphones1"
----
>>> "In summary, TinyGPT-V is a step towards making powerful AI models more accessible and efficient, which could lead to their use in a wide range of real-world applications1. The authors have also shared their code and training weights for others to use and learn from1"
-----
This is really interesting if you fan out implications over N time?
Here is my thinking:
Assume this paper results in a way of "compression-alyzed vision" into a model (a tiny compressed view into a model)
Then one, in a few years can imagine "laser views" - that slice through fractals of models to find the result. Resulting in tiny agents that have a heat-seeking-fractal-laser that can navigate giant data based on a method of knowing instantaneously what to exclude (meaning the path is defined by the walls that you already know you do not want to hit, so your steps are always that which helps you forward)
--
Or am I stating something obvious to all you brainiacs?
(no shame, I like thinking out loud)
This is a neural net built by conjoining Phi-2 [the best available small LLM, if you'll pardon the contradiction in terms] with pre-trained vision models like BLIP or CLIP. Models are piles of weights[/parameters] that are generated by training on datasets.
Already work has shown that training a multi-model model from the start results in smaller, more effective model. If you want to know more, check out recent work from CVPR [a vision machine learning conference] from '23[0] and upcoming work for this year[1]
edit to add:
The work of MS researcher Chunyuan Li[2] is worth keeping an eye on, particularly recent work like LLaVA-Interactive[3], a multi-model multi-task AI system, might be what you're trying to describe with your laser/fractal view phrasing.
[0]https://www.youtube.com/@VLPTutorial
[1]https://arxiv.org/search/?query=cvpr+2024&searchtype=all&sou...
If you check the dependencies of TinyGPT-V you'll see that it does not depend on tinygrad but rather torch... https://github.com/DLYuanGod/TinyGPT-V/blob/main/environment...