Chameleon: Meta’s New Multi-Modal LLM
arxiv.org
arxiv.org
> There's a Twitter thread from one of the authors ([2]). This part seems pretty important: "The models in this paper were done training 5 months ago. We've progressed significantly since then."
1. https://www.reddit.com/r/LocalLLaMA/comments/1ctsala/newly_p...
They also noted the problem was most pronounced once they got up to 34b sized. It’s a good reminder training large scale models leads to new interesting problems. I imagine a lot of techniques and know how are not published; all those little bits of experience add up to a lot of competitive advantage in many venues, so once again, thanks to Zuck and co for publishing.
I believe Meta is going to this direction with multimodality as well. The new GPT voice mode is probably using the same architecture.
What's mind boggling is that models perform better at the same parameter size with new modality added to them!
It seems obvious that 3D is the next modality.
Training time was 4282407hrs. At, conservatively, 200w gpus, that's (4282407*200)/1_000_000_000 GWh ~= 1 GWh. At 10c/kWh that's $100,000 ?
So if you have a single eqv GPU at home, it's 500yrs of training time and $100k in energy costs. Or, in practice, 3000 gpus for 2mo.
The AI industry has to hope the world doesnt change fast enough for these models to be useless.
EDIT: price is $100k
Edit: I see the price was updated
That $100k is conservative too, it doesn't include the cost of buying/renting the hardware, or the compute time spent on experimental training runs, or the cost of data acquisition, labeling and cleaning, or the cost of RLHF fine-tuning.
I would add “in their current form” and agree. There’s three things that can change here: 1. Moore’s law: The worldwide economy is built around the steady progression of cheaper compute. Give it 36 months and your problem becomes a $25,000 problem. 2. Quantization and smaller models: There’ll likely become specializations of the various models (is this the beginning of the “Monolith vs Microservices” debate? 3. E2E Training isn’t for everyone: Finetunes and Alignment are more important than an end to end training run, IF we can coerce the behaviors we want into the models by finetuning them. That along with quantized models (imho) unlocked vision models which are now in the “plateau of producivity” in the gartner hype cycle compared to a few years ago.
So as an example today, I can grab a backbone and pretrained weights for an object detector, and with relatively little data (from a few lines to a few 10’s of lines of code, and 50 to 500 images) and relatively little wall clock time and energy (say 5 to 15 minutes) on a PC, I can create a customized object detector that can detect -my- specific objects pretty well. I might need to revise it a few times, but it’ll work pretty well.
Why would we not see the same sort of progression with transformer architectures? It hinges on someone creating the model weights for the “greater good,” or us figuring out how to do distributed training for open source in a “seti@home” style (long live the blockchain, anyone?).
Now, I don't know of any distributed training technique that will make a significant impact on improving a model, and that security component is a big "if". But if something promising comes a long, I'd bet lots of people would be willing to donate some GPU time, especially if it were easy to set up.
I envision a future where crypto-style booms happen over tokens useful for purchasing priority computational time, which is earned by providing said computational time. This way researchers can daisy-chain their independent smaller rigs together into something with gargantuan capabilities.
On the plus side there's a lot of interesting things and it is generally easy to follow/figure out what they did.
On the minus side it's a little exhausting, and there's so much money in it feels like the vast majority of it is grifting. To add to that, the people who are trying to catalog it (the AI guy on LI) are the griftiest of them all.
I've found the best way to keep up is find one topic you want to learn about, deep-dive, and read all the related papers, then explore breadth-first from there until you find another topic...
It is all grifting. The moment someone creates something that can improve upon itself there will be an intelligence explosion and it won’t need press releases or debates about its intelligence. The current path of research will not lead to that, if there was something to be discovered here it would have been discovered already. It’s just new to the general consumer and there’s a wow factor associated with it like crypto and NFTs before. The truth is tech has lost its momentum and is desperate to find a new trick.
https://medium.com/@VitalikButerin/quadratic-arithmetic-prog...
Will read!
Is this accurate? I thought for example gemini pro used image tokens and gpt4-o similar
> without the need for separate image/text encoders
but then they say they pre-trained two different tokenizers, so maybe they just mean that the tokens go into the same attention layer? But then I thought that is how all the multi-modal stuff was happening already?
two typos stabilitize and multiplicate
Other work allows the model during training to learn the 'tokenization' more explicitly. that's more similar to Adept's Fuyu architecture, which I am personally a fan of, but also does not enable generating images out.
You can generate images using late fusion as well, though I am not aware of other public work that discloses both early fusion and image generation.
I understand why a unified model is an interesting thing to work on but doesn't the discovery of "modal-competition" suggest that at least short term it might be even better to train specialized models for each modality and some sort of modality-supervisor (glue code model)?
https://www.reddit.com/r/ChatGPT/comments/1bfa7s3/openai_cto...