Gato – A Generalist Agent
arxiv.org
arxiv.org
the point of this paper isn't "here, we solved general intelligence". It's "look, multi modal token prediction is a sound iteration". Look at the scale of the model in comparison to, say, gpt-3: this is a PoC, they didn't bother scaling it, because we've already seen where scaling these mechanisms leads.
What I would love to know is what kind of architectures deepmind et al are playing with in-house. Token prediction is a promising avenue, but it's more of a language that an intelligent agent may operate in, opposed to the self-sufficient structure of the intelligent agent itself -- the symbolic system that implements algos like gato. If that symbolic system will be the result of a generator-function, that generator function won't be token prediction by trade. I mean, maybe somewhere in the deep depths of a multi modal model, intelligent structure may emerge, but that would be a very weird byproduct.
yes, this kind of functional intelligence seems distinct from an actual living entity, which is the thing that uses subordinate functions to pursue goals and has some interior state, motivations and some sort of architecture. To reduce intelligence to tokens predicting more tokens is kind of like saying f(x), just solve for intelligence. When prediction itself is only partially what intelligent systems are about.
Agent is a very important word because it's accurate ("a means or instrument by which a guiding intelligence achieves a result") And it's the latter I think we ought to be after when talking about 'general ai'.
Scientifically this is unsatisfying but also if for some reason this turns out to be an engineering dead-end we have a big hole where a concrete theory of intelligence should be, with its components, mechanisms and so forth. And sadly I think this is still the weakest link in AI.
To me it seems a little bit like if you trained architects instead of having a theoretical basis for architecture, you just showed them every building in existence and sent them to work. It may very well work, but if it didn't you have a problem. And even if it did, you'd still want to have an understanding of why it works.
There are a lot of problems that do not have analytical solutions, like the N-body problem. With these problems all you can do is numerical simulation, which is pretty much what modern machine learning is.
On the other hand, for many reasons (energy, training/inference deployment complexity, latency, even some sense of model elegance I suppose), we don’t want to massively increase the number of parameters unnecessarily. But again, I don’t think we have great methods to estimate the “ideal” minimum number of parameters for a model to achieve its goals. And what we keep finding is that if you increase the number of parameters, and you increase the training corpus, the model gets more accurate, more impressive, and it’s not stopping.
So while I definitely agree that size for size’s sake is wasteful, I also don’t think we necessarily even know how to define “wasteful” for things like large language models right now.
In the case of GPT-3, scaling seemed to continuously improve results, they just kinda ran out of data. Are you implying this must be the same for this model? Or were you intending to say something different that I didn't see?
It doesn't have the active working memory to do so.
Diced onions > existential dread.
Although I do think that phrasing for the implication is already heavy handed in a negative bias.
I find it incredibly exciting; and it deserves more open discussion without provoking anxiety.
Right now, the state of the art AlphaZero models can destroy humans at Go. But what if the machine learning models could teach us things about how Go works that humans have not yet discovered.
Think about it like this: if the domain of a problem we want AI to solve is so complex that we can barely formulate the question, how could we be confident that we can understand 100% of the answer we get? “Here, gpu, make sense of this 20-dimensional problem my brain can’t even approximately visualize!”
But I would be satisfied if it developed a hypothesis in any language (including mathematics), not necessarily natural language.
Maybe it discovered some new science to pull that off?
But still, this is not cleverness, this just show that raw bruteforce + a few tricks can solve a few problems, by generating proofs of multiple terabytes(yes this is absurd scaling). The asymmetry between compute power and computer lack of intelligence is remarkable.
Why is this approach better?
The lack of fanfare on this achievement is baffling.
With "intelligence₁" we remain referring to "the ability to reflect on world representations achieving (recursively) refined concepts and selectively producing founded conclusions". It seems that the research relevant to this submissions is still alien to the context of AGI.
A claim that "this is the Dawn of It" requires clarifications and demonstrations (meaning, explicitness in clarifying why).
--
¹Some argument that it goes in that direction is in the Introduction: «Historically, generic models that are better at leveraging computation have also tended to overtake more specialized domain-specific approaches (Sutton, 2019), eventually».