Sequential modeling enables scalable learning for large vision models
yutongbai.com
yutongbai.com
It's only a matter of time before we have robots powered by large models pretrained to "predict the next token" across a bunch of different sensory modalities -- sight, sound, smell, touch, taste, etc. in a variety of artificial and natural settings, including social-interaction settings. Learning to read, learning to talk, learning to interact with the physical world, and so on -- all of it could very well be built upon the simple idea of learning to "predict the next token."
We live in interesting times.
this is all a bunch of gobbledygook until it isn’t.
The multi token[1] project which allows you to take any type of data and turn it into a token it's pretty interesting and seems like it's going in this direction.
I would really like to see a framework where you can take any modality of any type turn it into a series of tokens and just cram it into a language model and effectively turning into a multimodal model with almost no effort.
[0]https://general-pattern-machines.github.io/ [1] https://github.com/sshh12/multi_token
What would a Large Language Model that can manipulate audio-visual data as expertly as it can manipulate text look like ? This is beyond just Text to Speech or Captioning and Image Q&A. I think we'll find out very soon.
I think we might have audio2audio editing on the level of Stable Diffusion within a year or two. Based on recent progress with AudioLM etc.
For short clips, we will probably get video2video editing in the same timeframe. While it is computationally more challenging, it might end up being here before audio editing - because video is so popular in todays social media and marketing landscape.
Usable joint video and audio editing (that is coherent) will probably take longer.
In any case, I'm not sure Gato qualifies as a "large" model with 1.2B parameters -- it's kinda right below the threshold at which it could or would start exhibiting emergent behaviors. Maybe a new Gato with 10's or 100's of billions of parameters operating in the physical world?
I hope it didn't all get rolled up into Gemini and become a state secret they'll never publish on again, or lost in the shuffle in the chaos of the DeepMind/Brain merger/liquidation.
That's the most likely explanation, in my view.
See my other comment here:
I plan to buy a farm when I have the money and I'm pretty sure while I will/want to do a lot of hands on renovation and sculpting (park, etc) long term some type of robot should be good and affordable enough to take over when I'm too old.
Anyway I'm excited and looking forward to the code and models to be released, hopefully I can use them for my research! I think it's easy to overlook how revolutionary the transformer "way" of doing things has been, and the fact that so many different tasks can be reformulated in a "language" way I believe hints at something deeper about how the universe, our minds and language work.
Even though it makes so much sense, I never thought about it like this. Inpainting, Object Detection, Rotation, Lighting, Segmentation, Edge Detection, Pose Estimation, Surface Normal, Colorization and much more achieved by a single model.
I believe this and Codi-2(https://codi-2.github.io/) offer a glimpse of the future of Large Multimodal Models.
People will keep finding some small case or reason why not to call it AGI. And then finally once that last case is knocked down and we have agreement on a definition, we'll realize we crossed that threshold a "long" while back.
And I'm not saying we have AGI now, just that it's now clear to me how this process will play out.
(Where "long" in AI development timeliness probably doesn't mean the same thing "long" meant even in the 2010s.)
Invention preceding understanding is the norm, not the exception. We created fire before understanding chemistry, and we constantly use pharmaceuticals without really understanding how they work. Invention first, then theory comes along to explain and generalize.
For all we know, sentience is a necessary side-effect of semantic processing of any kind, in which case LLMs already have a form of sentience.
So yes, you're right that we don't understand our own sentience. In fact, we understand so little that it could literally be staring us in the face right now and we don't realize it.
Some people claimed that "creativity" (another rather nebulous term) was a uniquely human trait, and machines could perform tasks that require this. But recent generative AI models have started to make people question that position.
The visual prompting is a neat trick and it's great to see scale continue to work, but without comparisons to other models or releasing the code/weights, I don't think this strategy is going to be competitive.
It's relatively trivial to flip the sign on that training data and have a bot that will instead refuse to make antifascist or feminist arguments; if someone wants a Hitlerbot to ghostwrite "My struggle" for them, there's nothing that could prevent them from finetuning such a model from one of the publicly available models; there is no one that can enforce any 'moral standards' on the bots other than their creators.
I think we need a more rigorous definition to go beyond tautology
Your examples aren’t an issue I don’t think. To take one: efforts to minimize one’s consumption to combat global warming could fall in the “help tribe survival” bucket, which makes evolutionary sense as a base drive. And it seems like a reasonably “intelligent” response to that base drive — the predicted results seems positive.
• https://arxiv.org/pdf/2106.14742.pdf
• https://rmets.onlinelibrary.wiley.com/doi/full/10.1002/met.2...
• https://ieeexplore.ieee.org/document/9671442
Also non-transformer models, because of the scaling complexity on input: https://arxiv.org/pdf/2212.12794.pdf