Large language models can do jaw-dropping things. But nobody knows why.
technologyreview.com
technologyreview.com
> a hidden mathematical pattern in language that large language models somehow come to exploit... the fact that these things model language is probably one of the biggest discoveries in history
Is it? Excuse my ignorance. If we were talking about simple Markov chains, aren't these just stationary distributions?
For a given prompt/"state", an LLM essentially computes the next-state probabilities. This is done by compressing language and storing particular patterns/distributions. To date, we still don't understand what statistics are stored by LLMs. Or even what stats might be necessary to produce natural-sounding language. (A canonical Markov chain only works with n-gram statistics, but we know these are insufficient.)
IMO figuring this out to a human-level understanding could be a major breakthrough in science. It would reveal a deeper structure to language. It may even help understand how our brains process language; we'll at least know one plausible algorithm. Prior to LLMs, many scholars assumed language was unique to the brain and could only poorly be "computed".
"- You're just a robot, you can't write a poem
- Can you?"
It seems that GPT3.5 is already a much better writer than I am, and I'm still much better than a Markov chain, so LLM's are verifiably better.
Nobody wants to go near weeds. If ads were designed better we wouldn't be where we are now.
Bring back 10 second videos + skip - and relevant inline and non-manipulative ads and absolutely never endorse popups and we might have a fairer internet for content creators.
But here we are...
LLMs at least the most common ones can talk but they don't do things.
But even if they could. Humans can do it too, even your pet does jaw-dropping things. So far we never said that it's a problem in these cases.
So for me, such statements mostly communicate some sort of fear or skepticism. And I'm not saying that we shouldn't investigate why LLMs can do it. We should rather call it a research problem.
Similarly, LLMs have "hidden layer". We know everything about individual "neurons", but we couldn't connect them together ourselves in a way that training somehow does. The network that is the result of that training is "hidden" from our understanding even though its individual nodes are plain to see.
Since AI can increase in complexity while human brain can't, there may come a day when AI understands the human brain, without the human brain ever being able to understand AI.