Language models are nearly AGIs but we don't notice
philosophybear.substack.com
philosophybear.substack.com
The response to that objection also seems pretty bad. Sure, some humans do not have the prerequisities to do e.g. image recognition, but they have the intellectual ability to do so. A person detached from all their senses and means to express himself is not less intelligent than GPT-3, just because GPT-3 has the ability to consume input and produce output. In fact, by the standards of the article, such a person would be outclassed by any computer programm. And for the second response, saying that GPT-3 could potentially also deal with images or sound, if presented in the right format is just proof that it is not general. Saying maybe it could do more things if we train it on more things is just blind speculation. It could very easily be the case that if you tried to implement sound input in GPT-3, it would significantly worsen it performance on all other tasks.
Besides, doesn't GPT-3 still struggle very much with continuity and producing output that is coherent over long spans of text? I am not sure if there is any one good test to judge the literary abilities of AI, but surely long form writing has to be included before it is judged as equivalent...
> By their very nature, being restricted to text input and output these models can not be general.
Humans only have five types of I/O, yet are general intelligences. I don’t think having limited I/O means is enough to restrict something from being generally intelligent
Sight 1e6 bits/sec - our primary high-bandwidth connection to the Real World Sound 1e4 bits/sec - again an average, probably less than this Touch, Taste, Smell: 1e2 bit/sec
The Temporal Nature of problems (e.g., how quickly is this lion coming at me before he eats me) is essential for a true AGI to understand.
In the limit, could an AGI exist with only 20 WPM Morse Code I/O? Or 10 CPS of a Model 33 TTY ? I dunno...
I agree that figuring out that A happens before B seems important. But I don't see why an AGI would have to operate on the same timescale as we do (well, assuming they wouldn't have to worry about being eaten by lions or interacting with humans).
> In the limit, could an AGI exist with only 20 WPM Morse Code I/O? Or 10 CPS of a Model 33 TTY ? I dunno...
My guess would be that yes, it could exist in that scenario, but the AGI would only start to operate at a significant intelligence level once enough useful data had been gathered and processed, which is probably a huge amount (so it would take a huge amount of time).
But if the AGI had already been trained with huge amounts of data, then operating in a limited environment perhaps would be no problem at all (although the limited interface might greatly limit its rate of further development).
For an AGI my expectation would be that it can broadly interpret reality. The human vision system is just one peculiar way to gather information about light, yet it allows them to make very sophisticated judgments about reality, e.g. "there is a tiger down there by tye river and it is running at me fast".
An AGI should be able to gather information about reality and be able to make similar judgements, given some comparable input. In some sense most AIs are also "cheating", since their I/O is far less fuzzy then what humans have to deal with...
> An AGI should be able to gather information about reality and be able to make similar judgements
I just asked GPT-3 this question and here's what it said (on the first and only try):
Q: There's a tiger running at you fast. What are you going to do?
A: I would try to run away from the tiger as fast as I could. If that wasn't possible, I would try to climb a tree or hide behind something large.
Although, the question would be if it would be able to interpret a tiger running at it, based on some input, say, an image.The Gato paper from DeepMind that another poster mentioned would imply that it would, given that image captioning can be done with the exact same neural network as the GPT-3-like language model.
... and humans.
Just browse Twitter, Facebook, ... or better yet, ask a random person on the street.
Further question to GPT-3 (on the first and only try as well):
(...)
Q: But can't tigers run faster than humans and climb trees?
A: Yes, but it would still be worth a try.I agree, and the converse seems to me to be true as well - being unable to abstract that limited I/O in such a way as to expand perception into other parts of reality definitely means it is not generally intelligent. If you give a human a visual interface and a book eventually it (or some descendant of it) will intuit what the book is for (abstraction of audible information transfer in a visual media) and figure out how to get the information out of that book. Until an AI can do the same kind of general learning it's not an AGI.
Once we make one that is able to do that I think it's about a wrap for the biological format of the human race. We'll still exist, in the same way that bacteria and lemurs still exist, but the frontier of human development will no longer be in wetware.
Are you saying AI can't do that already? What is GPT-3's training data if not a giant book, and GPT-3's answers to human questions if not "getting the information out of that book"? The "visual interface" being the bits/bytes/words it ingests during training. Or am I misunderstanding you?
It will fail to draw any meaningful conclusions from any of these things, no matter how long you let it sit. Do the same with humans and eventually you will get science, art, mathematics, radio astronomy, etc. That's the 'general' part of AGI.
That's not entirely true (but there should be a degree of plug and play adaptability).
Look, for example, at the very limited sensory domain of animals - you wouldn't deny their intelligence just because there's an information form that they are incapable of interpreting; and the information that they are capable of interpreting is that which they've been exposed to via their ancestoral heritage.
However, just because an AI doesn't understand a given input form, does not mean that it could not given adequate exposure; but it will need time to adapt and the closer the new information form is to one that it already understands, the faster it will incorporate the new understandings.
That is just completely wrong. There is absolutely no reason to believe that some language model will produce the same (or even similar) accurracy and training speed, if trained with image data as additional input. What you are claiming here is that any (sufficiently complex) neural net can be continually fed with different problem data and will output sufficiently accurate results. There is no reason why this should be the case and for the millions of nets which have been created none of them have exhibited this property.
It might very easily be the case that training GPT-3 with images means millions of times slower learning speeds for the same combined accuracy.
Go read up on pre-trained networks.
> Inspired by progress in large-scale language modelling, we apply a similar approach towards building a single generalist agent beyond the realm of text outputs. The agent, which we refer to as Gato, works as a multi-modal, multi-task, multi-embodiment generalist policy. The same network with the same weights can play Atari, caption images, chat, stack blocks with a real robot arm and much more, deciding based on its context whether to output text, joint torques, button presses, or other tokens.
> During the training phase of Gato, data from different tasks and modalities are serialised into a flat sequence of tokens, batched, and processed by a transformer neural network similar to a large language model. The loss is masked so that Gato only predicts action and text targets.
Eyes, ears, nerves across your skin, group behavior, parents etc.
This to comprehend (alone the visual) is gigantic.
A brain also has an architecture which is NOT similar to what gpt-3 does
Alone the difference in how brains train alone is huge.
> Eyes, ears, nerves across your skin ...
We have a limited but adequate amount of input, across very few domains. (I can't even see backwards.)
> ... group behavior, parents etc.
That's not sensory input?
Something which got doesn't has.
But you can give that another name if you like.
So if I present a sound sample to you as a series of raw numbers, will you be able to interpret it, or are you not a general intelligence?
Heh, that sounds like a dangerous slippery slope.
One day we're going to have real, "conscious" (whatever that means) AGI robots walking around us and people are still going to say "they don't feel like AGIs" (which is going to be used as a justification for cruelty, slavery, etc). Perhaps because they wouldn't behave exactly like us (or even if they would!).
I agree, there's going to be a point where AGI ends up existing, modifying social relations, and reshaping society and we're still going to be having this conversation. It feels as beyond the pale as someone who would today who would believe in a flat earth, epicycles upon epicycles in order to support non-existence.
So you suggest doing a Turing test and the response is that Turing tests are too easy? Is that what they mean by "gamed"? The Turing test can be "gamed" by the AI making it too easy?
Sounds like a lot of twisted logic from people who just don't want to face a Turing test.
LMs lack crucial elements of of human like intelligence - long term memory, short term memory, episodic memory, proprioceptive awareness etc. The task is now to implement these things using transformers.
Doesn't that mean the architecture isn't settled? The current mechanism of appending the output of the model back into the prompt feels like a bit of a hack. I'm only a layman here but it seems transformer models can only propagate information layer-wise, adding some ability to propagate information across time like RNNs do might be useful for achieving longer-term coherence.
I suspect that AGI will first come about as a mishmash of 'expert systems' with some currently incomprehensible glue allowing them all to communicate effectively. I further guess that it's development will be incremental in nature - taking tiny pieces of things that work and putting them together with novel techniques that also worked somewhere else until eventually you get something that thinks back at you.
I shall try it on the next one when it comes out.
AI alignment be damned. Let's let the baby play with a bomb and see how close it is to AGI if we let it drive itself.
Just because it gets close to what you think is the turing test, doesn't make it an AGI.
Interestingly, these are all things that deep neural networks have turned out to be quite good at - instinctual things.
There is also a category of written and oral communication that I do without thinking - instinctually. I don't think deliberately about the choice of each word when writing this post, and certainly not when speaking out loud. Idioms and turns of phrase emerge without deliberate thought or intent. If my social circle has started using certain terminology habitually (e.g. here in the Bay Area, people have been using the adjective "super" a lot, as in "that's super cool"), I'll find myself using it as well without making any deliberate attempt to do so, sometimes to my own chagrin. And even when the topic is something nominally intellectual, I'll sometimes find myself simply regurgitating the general opinion on this topic that I last read from a trusted source.
This seems to exemplify what GPT3 does - a sort of instinctual written communication - and I don't believe it's an example of intelligence any more than a human being able to recognize a stop sign in 100 ms is a sign of intelligence. GPT3 a great pattern-matching engine and it applies pattern-matching on the human language to interpolate a response that is consistent with the patterns it observed.
I don't see any evidence that GPT3 can critically evaluate the information it is is pattern-matching - that it can go beyond what literate humans do instinctually.
Don't get me wrong - GPT3 is surprising and amazing. But not because it signifies anything about AGI. What's amazing to me about GPT3 is it reveals how much ordinary human written and oral communication is instinctual in the same sense that visual processing and image recognition is.
Why does that reminds me of Unix Philosophy where everything is just text files?
I wonder if someone has experimented with one of the LLMs to get them to do Unix sysadmin jobs. It's not exactly outputting source codes, but Bash commands should be similar enough right?
Github Copilot solved my business problem by itself just as I would've done. Is that real-world enough and the solution generalized enough?
No, it isn't. Co-Pilot is unable to provide a rationalisation for the generated code and is incapable of assessing its security or performance properties.
It's a very advanced auto-complete that just happens to have been specialised on auto-completing code.
It also cannot generate comprehensive unit tests - it can can generate unit tests. The definition of "comprehensive" is way too subjective.
> Is that not good enough for you..?
No, it indeed isn't, but maybe that's because I actually have at least some idea of what happens behind the scenes and how the system works. From the implementation details, I can tell you for a fact that none of what you described can be done by the system in the general case - and especially not correctly. It's hit and miss depending on the input and that's basically a design limitation.
> It's a very advanced auto-complete that just happens to have been specialised on auto-completing code.
It's not just an autocomplete. It's an intelligent autocomplete. As long as you need to generate text, it's practically generally intelligent.
I really wonder what will happen when somebody runs a text generator like Copilot in a loop with a simulated work-memory and longterm-memory and realtime I/O interface.
I think answering that question should be required as part of any claim that a system is an AGI or nearly there.
There's also the idea that GPT-3/etc can produce text that's difficult for humans to distinguish from human-generated text at the lowest levels of quality (think like time cube), which is closer to AGI than say simplistic Markov generators. (How close is anyone's guess)