John Carmack on the Similarity of Human Learning and LLMs Training
twitter.com
twitter.com
LLMs are exposed to human-curated data. Let's see an LLM curate its own data out of nothing but experience of raw continuous signals.
Sure, GPT trained on the internet is a nice way to condense and retrieve human-curated data. But it's not going to give us a Lt. Data or C-3P0 that can learn and adapt in real time.
Yes humans need to learn to extract words from signals, but that does not mean it's not correct to count the number of words they've been exposed to in their lifetime. Your comment is like saying we can't count the number of hamburgers someone has eaten, because they actually eat myofibrillar proteins, not burgers.
https://miro.medium.com/v2/resize:fit:1056/0*E1eNateTiDThGcY...
Buckle up.
I see the progress of AI stagnating while people board the next equivalent of Crypto craze because someone want to financially profit off something not fully seeing realization.
However it is not without merit which Crypto was completely without merit and something to show for itself.
As flawed and imperfect ChatGPT is it clearly mirrors its creators and that is a compliment.
Humans learn based on human curated data too.
School is not natural. Our whole env is neither. Babies can't survive.
I would even go so far to say that the potential model a LLM would create internally might not be that far away of that of a human.
And segment anything was just announced. The performance of zero shot systems is tremendous.
It's not far fetched to assume that chatgpt combined with segment anything together would allow it to create an even more accurate model of the world.
> School is not natural. Our whole env is neither. Babies can't survive
A human goes to school to learn.
Humans in general learned based on experience of the world around them. We invented language, no one taught it to us. We learned to make fire, forge tools, cook food, practice medicine, etc on our own.
School is just how we pass down that learning.
Individual humans were taught those things, either directly or indirectly by observing / listening to others.
We do learn through our experience, of course. But most of what we learn is from others in one form or another.
A lot of these children had trouble with learning language after they returned to human civilization.
I am of the opinion that we'll know when we have real AGI because it'll be able to fold my laundry. One can dream.
On the other hand, I'm not so sure an LLM couldn't invent a language, particularly one to use with another AI. It's sci-fi from 1970, but I just rewatched "Colossus: The Forbin Project"[0] recently which features this as a plot point. Really enjoyed it years ago when I first saw it, but it's improved with age (at least for me) now that we are closer to AI.
Plenty of studies showing long term academic achievement differences based on number of unique words babies are exposed to.
And I promise you, having a baby/toddler is really damn close to doing data labeling. Reading picture books, you are basically labeling objects. Walking around the grocery store, you are labeling objects. All the time, again and again, and the same object will get labeled repeatedly, and sometimes it will even be fact checked. A toddler will point to something that they know is an orange, ask "apple?" and you had sure as hell better reply "orange".
It's analogous, sure, but not really in the same ballpark. There is no AI system today that you can teach by pointing to things and making noises with your mouth. Why not? What's missing? An awful lot.
We've got gesture recognition, so we can do pointing; and that part of the system taken in isolation is not made any harder when the label for an image happens to be from, say, a latent space vector from the interior of an unrelated GAN whose sole purpose is to turn audio input into the specific set of muscle outputs which cause the vocal tract to produce output that sounds the same as a recently heard noise.
Of course, one could easily build a model whose tokens line up with qualia we experience.
Nobody really knows what is necessary or (let alone and) sufficient for having a what-is-it-like-to-be-ness.
Lots of us claim we know, but whenever I've asked for details, it's always been some combination of
(α) circular, e.g. "consciousness is sentience, sentience is consciousness";
(β) something that's so vague it accidentally includes VCRs because they can record information that can be recalled later;
and/or (γ) that doesn't include all humans (most often by excluding those with aphantasia and/or amnesia, but not only).
So… perhaps any given LLMs does, and perhaps it doesn't. Perhaps the structure can support qualia, but only actually gets them when the training process passes the analog of a thermodynamic phase change; or perhaps, even if that's the right general idea, qualia is a different phase change that this model can't support.
Unfortunately, I can't find that phase change analogy with a quick web search, so I'm not sure if that analogy has ever really been made before or if it was just a dream I've misfiled as a memory of reality.
Personally, I’m inclined to believe that consciousness requires an integrated attention-self-world model, but maybe not much more that that.
Thanks!
https://the-decoder.com/john-carmacks-general-artificial-int...
It still seems to be missing any sense of what is True, not sure if that’s possible to embed in that model or if we’ll just have some human feedback hacks and eventually get a better model
I feel like putting the word 'forever' ruins the point. It's the most extreme strawman.
It's amazing how the scaling has unlocked the emergent behaviors! When I look at the scaling graphs, I see that the ability to reduce 'perplexity' is continuing with scale and capital investment with no sign of slowing yet (it will slow eventually). I also see that reducing perplexity is continually unlocking new emergent behaviors. So I would guess that scaling will probably unlock so many more new emergent behaviors before it eventually plateaus!
While I definitely get "longer" and slightly more "in-depth" answers from GPT-4 vs 3, it already feels like that capability growth curve is starting to plateau.
I strongly disagree. Anyone who wants to look for themself can see the GPT 4 technical report.
The capability curve will necessarily appear to plateau when the starting point is as good as it is now. The improvements we recognize will be subtler. Halving the remaining error rate will look less impressive for each step.
ChatGPT is quite good at something like "putting together facts using logic". Things medical diagnosis or legal argument. However, if the activity is "reconciling a summary with details", ChatGPT is pretty reliably terrible. My general recipe is "ask for a summary of a work of fiction, then ask about the relationship of the summary to details you know in the work." The thing reliably spits out falsehood in this situation.
If it could reliably tell me when it DOESN'T know something, I'd have a lot more respect for it's capabilities. As it stands today, I'd feel I need to fact check nearly anything it gave me if I'm in an environment that requires high levels of factual accuracy.
Edit: To be clear, I mean it telling me it doesn't know something BEFORE hallucinating something incorrect and being caught out on it by me. It will admit that it lied, AFTER being caught, but it will never (in my experience) state that it doesn't have an answer for something upfront, and will instead default to hallucinating.
Also - even when it does admit to lying, it will often then correct itself with an equally convincing, but often just as untrue "correction" to its original lie. Honestly, anyone who wants to learn how to gaslight people just needs to spend a decent amount of time around GPT-4.
[To be clear, I did not tell it it was from after the cutoff]
Funnily enough, the Geoffrey Hinton of today is probably some symbols researcher, shouting in the desert, like Hinton shouted in the 1980s, that we need more than matmul.
One interesting, biology-inspired mechanism, would be quorum sensing [1]: a basal cognition-like, decision-making function in which decentralized systems (bacteria, cells) start building functionality (sensing/decision) from the bottom up. There is no hint for this sort of mechanism in our current artificial 'neural networks'. Not that it should, but our cells and in general cells use these kind of 'tricks' to solve problems in all kinds of spaces (transcriptomics, morphogenetics, etc.) without requiring ridiculous amounts of energy, time, or other resources.
Now, as far as whether LLM or offshoots can get to AGI, I know plenty of good arguments for them not being able to do that. I don't think the claim that they're just accumulating more abilities in each iteration in an inexplicable way is true. But I've lived long enough to know that you should never get too cocky when one is "arguing with success". So maybe.
Moreover, given that LLM programming is basically just bucket chemistry, if an LLM can the ability to competently pursue long term goals, it seems like it will have a good chance of some of its goals being random cruft that will make it quite dangerous.
Equating LLMs to parrots is offensive, to the parrots. Well, maybe not parrots, but crows are intelligent: "Scientists [5] demonstrate that crows are capable of recursion—a key feature in grammar. Not everyone is convinced" [6].
[1] https://en.wikipedia.org/wiki/Hybrot
[2] https://en.wikipedia.org/wiki/Aladdin_(BlackRock)
[3] https://www.nytimes.com/2023/03/31/technology/sam-altman-ope...
[4] https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nomi...
[5] "Recursive sequence generation in crows", https://www.science.org/doi/10.1126/sciadv.abq3356
[6] https://www.scientificamerican.com/article/crows-perform-yet...
That said, I did get a chance to speak with him one on one last year and he really emphasized a few things; the need for being product oriented and giving customers what they want rather than chasing cool engineering (über)solutions; not being careless with resources just because we have more power with modern hardware (he poked fun at React where you spin up a new thread just for an interactive button); and being aware of the inefficiencies brought on by infinite resources (# of engineers and/or funding, which make you think less critically about timelines and delivering within bounded means)
He said games were one of the most complex things humans build and (with implied comparison) the mathematics and physics of rocketry hadn't changed much since the 1960s. Sounds true to me for the 2000s.
Where was the utter failure and first principles?
Wow, you think his ability is straight performance optimisation? A big part of his early fame was from doing research to find good algorithms to achieve his goals. He also had that stint in real time control systems... flying hovering rockets before spaceX even existed.
He's a problem solver with a strong ability to sift through possible solutions for what actually works, and quite capable of devising his own solutions when is research comes up empty.
But yeah, he can write assembler too.
That is what sets up their mental model of the world, and language is fit over the top of that. It is hard to assess neural net efficiency with that difference in place.
Vision doesn't seem relevant for what's missing in general intelligence: there are people blind from birth that nonetheless get very fluent in thought and language. It doesn't cause many significant developmental delays other than in some vision associated stuff like early motor skill acquisition, social stuff from not being able to see facial expressions, etc.
The models may be missing data grounding in 3d space that people born blind do have, like proprioception, touch, and hearing. Deaf and blind from birth can cause more significant developmental delays, but becoming deaf later after some critical threshold doesn't seem to, so you are still looking at potentially only a short amount of years worth of data that seems to be needed (Hellen Keller went deaf and blind at 19 months).
No other animal can learn human language or logic to the extent that humans can, no matter how much you train them. But humans learn language easily in the first few years of life.
It’s almost as if the human brain is preprogrammed with the general concepts of all languages, and it just needs to be fine tuned with specific vocabulary and grammar rules.
1-2 trillion tokens of training for gptx class models
It would take 20 years for a human to read that many tokens at 8 hrs of reading per day (although i'm pretty unclear on whether that's unique token sequences or randomly selected token pairs -- feels like a big difference between the two!!)
We are all here now -- who thinks we are anywhere of note together?
Personally I think that models that do not experience incremental change remain an entirely separate class of intelligence from those which evolve continuously -- but I am open to being persuaded otherwise (certainly I never thought making gpt2 bigger would be as impressive as gpt4) ...
>It would take 20 years for a human to read that many tokens at 8 hrs of reading per day
Normal reading is ~200 words per minute.
That's 12,000 words per hour.
That's 288,000 words per day (reading for 24 hours straight).
That's 105,120,000 words per year.
That's 2,102,400,000 words per 20 years.
A trillion is 1,000,000,000,000.
I think people overestimate the data efficiency of the biological neural networks w.r.t. transformer models.
Two obvious differences are that humans also learn from the context in which they live (social, environmental) and humans have to basically learn how their own wiring works. And their wiring is/can be highly variable.
> more like a thousand times more. > Between 1 and 2 trillion tokens. > It would take a person 22,000 years to read through 1 trillion words at normal speed for 8 hours a day.
It's kind of surprising to me that around 1000 humans could feasibly read the whole internet (as scraped for LLMs) between each other in only 22 years. The internet feels so much more massive than that. Though of course lots of the stuff on arxiv etc. is so dense there is no way you go through it at normal reading speed.
He's not the world's leading researcher or anything, but he's far from green in this space.
edit: oh, does he mean a billion total rather than a billion unique words? haha i r smat
Then add in the fact that we are only talking about language acquisition and all indications are that sufficiently large models (And the human brain is quite the large model) have significant benefits from cross training and it starts to become clear how false of a premise this is. 11 million bits of neural information are being pumped to our brains every second of our lives. That’s 80 gb of external data ingested during our waking hours in a single day plus who knows what sorts of internal synthetically generated data, all of this just for fine-tuning.