That's how imagenet works. The curves and lines that compose things are organized hierarchically into bigger and bigger concepts hooked together with weights. Language models just do the same with written concepts. When you have enormous models it just goes beyond all human intelligence to contain it as a subset of our understanding, and in ways it is smarter than humans, because we can't grasp its enormity.
However, you might not be aware of the current research into "Emergent Phenomena in Large Language Models"[1].
Work such as PaLM[2] show that language models evolve the ability to "reason" over the internal data representation. Minerva explicitly breaks down that ability by "comb[ining] several techniques, including few-shot prompting, chain of thought or scratchpad prompting, and majority voting, to achieve state-of-the-art performance on STEM reasoning tasks."
This lets it solve novel problems such as "Assume that the variance of the first n natural numbers is 10, and the variance of the first m even natural numbers is 16. Compute m + n".
To preempt the "oh it is in the training set" issues, I'd note this (on a different problem): The model correctly takes the derivative of a position vector to compute the velocity. It then computes the norm of the resulting vector, and uses a trigonometric identity to obtain a final numerical answer. Minerva takes a completely different approach from the ground truth solution."
I'd encourage you to read the Minerva paper[4] and update your understanding on what large neural networks are doing. These reasoning abilities are not* as simply as "compressed data representation + reconstruction" - the chain-of-reasoning approaches partially use that of course, but also appear to a form of lemma chaining, similar what something like Prolog does.
As a specific demonstration for their ability to generate novel outputs that make sense:
Prompt: explain Algorithm X in the style of a Kayne rap
Response: Alright, here's a rap about Algorithm X[5], the famous algorithm for solving the exact cover problem:
`Alright, here's a rap about Algorithm X, the famous algorithm for solving the exact cover problem:
Verse 1:
I'm Algorithm X, the greatest of them all
I'll solve any exact cover problem, big or small
My steps are simple, my logic is sound
Just give me your matrix, and I'll give you the answer you found
Chorus:
Algorithm X, the master of the game
I'll solve your problem, and do it by name
Exact cover, no problem for me
I'm Algorithm X, the MVP
Verse 2:
First, I'll start by selecting a column
Then I'll look for a row with a one in that column
Once I find a match, I'll mark the row and column
And repeat the process until the matrix is gone
Chorus:
Algorithm X, the master of the game
I'll solve your problem, and do it by name
Exact cover, no problem for me
I'm Algorithm X, the MVP
Verse 3:
But sometimes the matrix is too big to solve
That's when I use my dancing moves
I'll shuffle and prune, until the solution is clear
And in the end, I'll have your answer, my dear
Chorus:
Algorithm X, the master of the game
I'll solve your problem, and do it by name
Exact cover, no problem for me
I'm Algorithm X, the MVP.`
Now I entirely concede that it has read things on how Algorithm X works, and that it seems to use a template for "rap" responses. But: But sometimes the matrix is too big to solve
That's when I use my dancing moves
I'll shuffle and prune, until the solution is clear
And in the end, I'll have your answer, my dear
I refuse to believe that anywhere, at any point has someone written an explanation of the use of dancing links[6] in Knuth's Algorithm X like that.[1] https://ai.googleblog.com/2022/11/characterizing-emergent-ph...
[2] https://arxiv.org/abs/2204.02311
[3] https://ai.googleblog.com/2022/06/minerva-solving-quantitati...
[4] https://arxiv.org/pdf/2206.14858.pdf
In an odd way, it kind of reminds me of the beginning of the Bible.
> In the beginning was the Word, and the Word was with God, and the Word was God.
Has a large language model feel to it, doesn't it? hah.
The leading labs (Google Brain/DeepMind/NVIDIA/Meta/Microsoft/OpenAI) all publish in the open.
I'm excited by three things:
This emergent phenomenon thing - as we build bigger models there is a step function where they suddenly develop new abilities. Unclear where that ends.
The work people are doing to move these abilities to smaller models
Multi-modal models. If you think this is impressive just wait until you can do the same but with images and text and video and sound and code all in the same model.
No. It's pretty clear what the state of the art is.
For the rest of your questions I'd say I do spend a lot of time thinking about what intelligence is.