AI research is almost as purely empirical as the gradient descent loops its practitioners use to optimize their models. “Why” anything at all works is barely an afterthought.
AI research is almost as purely empirical as the gradient descent loops its practitioners use to optimize their models. “Why” anything at all works is barely an afterthought.
That's a very different claim from being "poorly understood" though. The emergent properties of any system with billions of parameters is hard to understand completely, that's the fault of data science more than computer science or even mathematics.
Understanding does have layers, and that's why "poorly understood" is a meaningless goalpost. A book can be well understood without researching the gematria behind character's the names when you write them in reverse. An LLM can be well-understood even if you don't comprehensively test each quantization for miraculous unexpected behavior at the FFN level.
Notably, you could still print them all day.
I think we actually all know what "poorly understood" means. There's no need to play tedious semantic games.
All that is to say, being able to build something is not not not the same thing as understanding it.
Because the extent to which we don't understand the brain, is quite overpowering.
Some people forget that when they say "but it's not different from what a human does" ...
We know lots about human development and genetics and biology and evolution and neuroscience and the physics of how brains are connected and send signals and how generally they are put together and have names for their parts and all that, but we’re clueless when it comes to “the hard question” of how qualia and consciousness emerges from that.
The scenario with the spooky simulation of thinking that emerges from LLMs is in the same category, with different details. Lots of knowledge about the substrate of the phenomenon, little to none about the much bigger question of how we get the appearance of cognition from these trained artifacts.
Clearly we understand extremely well how LLMs are created mechanically. We invented them and are currently putting massive amounts of work into studying and improving them. But that work is perforce largely empirical; figuring out the why once again eludes us. It just goes to show how mysterious the underlying phenomenon of cognition is.
Are you saying that thinking and cognition requires language use? Cause I think not.
Or are you saying that language use is sufficient for cognition and thinking? Cause I'm also not convinced of that.
What I am convinced of, is that a machine capable of language use is capable of tricking people into believing there's a "there", there. In pretty much the same way as the famous supra-normal stimuli experiment made baby seagulls believe that a stick with a red dot was their parent. It's exploiting our instincts.
Well-informed people understand that LLMs work as well as they do for the same reasons as horoscopes, fortune-telling and homeopathy.
Are you suggesting that gradient descent is an empirically found and not understood technique? It was originally proposed by Cauchy in 1847, its properties are very well understood.
You might be referring to properties of the domains its being applied to.
The data is the input, the output is to generally find the lowest amount of a loss function. It’s a greedy approach because brute forcing is inefficient.
It’s no more empirical than a greedy algorithm for scheduling.
Right, GP is drawing a distinction between search, ie mechanical exploration of a space, with understanding, ie having a map of the territory such that you don’t need trial and error.
"Empiricism" implies that the technique is based on observable, but not mathematically proven foundations. If a problem space is convex, gradient descent is guaranteed to converge to a global optimal solution, regardless of whether you know the exact formulation of the space.
Applying it when you don't understand if a space is convex is another question, but that's not a fault of gradient descent.