Evidence of a predictive coding hierarchy in the human brain listening to speech
nature.com
nature.com
As you listen to someone, your brain is constantly matching the sounds arriving at your ears with a prediction of what the next few words might be. Listening in a non-native language, my predictions about what comes next aren’t very well tuned at all, so if I can’t hear every word clearly then I can easily get lost.
Another signpost: sometimes you mishear someone — “oh, I thought you said xyz” — but the thing you thought you heard them say is never gibberish, it’s a grammatically and contextually valid way to complete the sentence.
Language models are just missing some component that we have. The method for deciding what to output is wrong. People aren’t just guessing the next sound. It’s like they said, there’s multiple levels of thought and prediction going on.
It needs some sort of scratch pad where it keeps track of states/goals. “I’m writing a book” “I want to make this character scary”
Currently it only works on the next tokens and its context is the entire text so far, but that’s not accurate. I’m not deciding what to say based exactly on the entire text so far, I’m feature extracting and then using those features as context.
e.g She looks sad but she’s saying she is fine and it’s to do with death because my memory says her dad died recently so the key features to use for generation are: her being sad, her dad died, she may not want to talk about it
Fortunately I stopped only one syllable into "birthday".
It didn’t matter what it was for. Birthday? Best wishes, Dave. New born? Condolences on your loss? Retirement? Best wishes, Dave.
His name was Albert.
Regardless, I find myself as I age unsure of how to respond to people and default to similar behavior. Unplanned pregnancy? Is that a surprise miracle or an unwanted interloper? “Wow you’re in for an adventure, I’m happy to support you”.
This is why I like to surprise my friends, family and coworkers with the occasional curveball
> Maybe articles on AI just trigger me and I spew the same arguments. I
But you are not representative of all humans.
> Doesn't make it not mecahnical.
There are mechanical things that are more than just prediction machines. Why did you make the "leap non-LLM" == "not mechanical"?
I actually started to type almost the same reply as your parent earlier, but did not post it. I used "difference in quantity, not quality" instead of "scale", but I also included the self observation. So maybe that makes two of us.
This concept is literally known as a "grandmother neuron", and its widely considered to be debunked. I refer you to this neuroscience lecture.
https://www.youtube.com/watch?v=_njf8jwEGRo
Beyond the simplest, most common tasks where it pays to have a rote lookup table, you truly do need an algorithm.
We have essentially built a computer-based replica of our internal language engine.
The goal was always to mirror ourselves.
So the other side of the above statement could be: “oh wow, LLMs are just like us”.
We should be very impressed instead of dismissive.
At the same time, maybe yes, our language capabilities can be completely imitated by LLMs.
What do we do now?
How can we honestly say that’s the case when we cannot explain (or define) our “internal language engine”?
We cannot explain other people, but we imitate them.
Kids learn a lot of stuff by imitating. It’s a basic human function.
So essentially we created systems that we cannot fully explain, to imitate our own systems that we can’t fully explain either.
>Me: Miserable. I'd like an Iced Latte.
...
>Barista: what type of milk?
>Me: Straight out the cow.
>Me: 7/8ths.
>Me: Chocolate.
Maybe I'm just a high entropy individual.
To trivial questions no, but to more complex questions humans actually does go "hmm, let me think". ChatGPT doesn't do that, it just blurts out the first thing that gets into its head regardless if the question is trivial or extremely complex.
For most of our "hmm, let me think"'s I'm not sure what we do is significantly more complicated than that. We just get to hide the inner monologue.
Maybe these language models could be way smaller and cheaper if we added a recurse symbol to it that made it iterate many times. Hard tasks would still take a lot, but most banal conversations would be very cheap.
I'm not sure I know what you're talking about.
https://www.reddit.com/r/AskReddit/comments/ajhnv/sit_or_sta...
I misremembered, it was wiping apparently.
I love your enthusiasm in this direction (I really do).
But here's some free advice I won't be able to prove for a long time:
Anyone who is currently convinced that somehow our fledgling efforts in the direction of building useful ML models are somehow going to yield a new golden age of neuroscience and understanding of the brain-- rather than the other way around-- is gonna be in for a long and frustrating next couple of decades, especially if they're low-openness types.
Eh, HN is supposed to be "better" than your standard public forum and we still always end up making habitual arguments, not truly reasonable ones. I find your hope inspiring but not enough to ignore my senses.
It can if you give it the option. My own openai chatbot prompts to either respond or to ponder by stating a question or considering a related idea. It infrequently will decide to ponder for one to about a dozen times as it restates an idea to itself in various forms, which gets recorded into the growing conversational prompt.
Elsewhere in the thread, someone mentions it always uses the same amount of time. In my estimation, it will spend longer on introspective or recursive prompts. Easier to get it ranting absurdities using those as well.
It's always the craziest chatter when the request takes a couple minutes to get back to me.
Above them would be creatures living in a discrete and almost continuous world of rational numbers. They would have highly sophisticated and elegant art, and their science would almost always get close to truth, but never touch it - the limitation of rational world.
Yet above them would be the god-like creatures inhabiting a world of continuous real numbers. They would seem a lot like the creatures right below them, but incomprehensibly greater in reality. They would look transcendent to the rational creatures.
Even higher would be the hyper-continuous worlds, but little would be known about them.
The question is where we are on this ladder.
And I believe we have the technology and advances we have because of this. Can you imagine if you had to devote actual brainpower to every inane thing you encountered in your day? You'd be completely exhausted within two hours of waking up. Every time my brain reflexively reacts to something based on past experience I'm thankful I didn't have to think about it. I can spend my finite energy on something interesting and novel.
My dad used to like putting sales people off their scripts by answering such questions rudely.
"How are you today sir"
"Do you really care?"
And I promise you they thought hard about their answer to that question too.
Maybe its just because you're focusing on the form like niceties questions and not real questions people deal with. Good Morning isn't even a question, per se.
Ai ai ai. This is bad, so bad. It's a classic case of p-hacking. They took some mapping of language model activations and moved it around a mapping of brain activity until they found an area where the two correlated- and weakly at that, only at a low R = 0.23.
Even worse. They chose GPT-2 over other models because it best fit their hypothesis:
For clarity, we first focused on the activations of the eighth layer of Generative Pre-trained Transformer 2 (GPT-2), a 12-layer causal deep neural network provided by HuggingFace2 because it best predicts brain activity7,8.
Not only the model- its activation layers.
They shot an arrow, then walked to the arrow and painted a target around it.
Dear god. That gets published in Nature? Phew.
The paper is from 2023 but their info is totally out of date. ChatGPT doesn't suffer from those inconsistencies as much as previous models.
ChatGPT (and indeed all recent LLMs) using much more complex training methods than simply 'next-word prediction'.
* one, applicable to current language models (which ChatGPT is one of them), claim that they "they fail to capture several syntactic constructs and semantics properties" and "their linguistic understanding is superficial". It gives an example, "they tend to incorrectly assign the verb to the subject in nested phrases like ‘the keys that the man holds ARE here", which is not the kind of mistake that ChatGPT makes.
* Another claim, is that "when text generation is optimized on next-word prediction only" then "deep language models generate bland, incoherent sequences or get stuck in repetitive loops". Only this second claim is relative to next-word prediction.
for the first two, I think this orthogonal
|-----long timing loop / top of parse tree-----|
| |
|-shorter / child node -| |-shorter / child node-|
| | | |
|highest freq| |highest freq| |highest freq| |highest freq|Everyone forms those predictions, it's how they come to an understanding of what was just said. You don't necessarily memorize just the words themselves. You derive conclusions from them, and therefore, while you are hearing them, you are deriving possible conclusions that will be confirmed or denied based on what you hear next.
I have an audio processing disorder, where I can clearly hear and memorize words, but sometimes I just won't understand them and will say "what?". But sometimes, before the other person can repeat anything, I'll have used my memory of those words to process them properly, and I'll give a response anyway.
A lot of people thought I just had a habit of saying "what?" for no reason. And this happens in tandem with tending to complete any sentences I can process in time...
What's it called? I do this sometimes also and I'd like to know more.
I think it's called "Auditory Processing Disorder"[0]. I'm pretty sure it has to do with me being autistic. I've done hearing tests before and my hearing is just fine, it's just processing what I hear that is the problem.
"Sometimes saying [“huh,” “what,” or “I don’t understand”] and then immediately responding appropriately"[1] is exactly what happens with me.
[0]: https://en.wikipedia.org/wiki/Auditory_processing_disorder
[1]: https://www.vocovision.com/resources/parents/auditory-proces...
Not sure if I'd qualify for a diagnosis, but either way, it's cool to have confirmed that such experiences are 'a thing'.
I also only noticed because people would ask me to stop saying it, and then I would immediately say it anyway because it wasn't compulsive and I really couldn't tell that I was right about to understand what they said. I hadn't yet figured out that there was just a delay sometimes.
It makes complete sense you would have an idea of the next word in any sentence and some brain machinery to make that happen.
It in no way means you’re just a LLM
"and some brain machinery to make that happen" - Getting close to not having a lot of "brain machinery" left that is still a mystery. Pretty soon we'll have to accept that we are just biological machines (albeit in the form of crap throwing monkeys), built on carbon instead of silicon, and we run a process that looks a lot like large scale neural nets, and we have same limitations, and how we respond to our environment is pre-determined.
If someone says "Computers will never be smarter than humans", then sure go ahead, post it. But most of the time it is just repeated whenever someone says that ChatGPT could be made smarter, or there is some class of problem it struggles with.
The objection people would come with then is something like "but we could add those other systems to an LLM, it is no different from a human!". But then the thesis would be "humans are no different from an LLM connected to a whole bunch of other systems", which is no different from saying "some parts of human thinking works like an LLM" as I suggested above.
Recently stuff like ChatGPT is challenged by people pointing out the nonsense it outputs, but it has no way of knowing whether either of its training input or its output is valid or not. I mean one could hack the prompt and make it spit out that fire is cold, but you and I know for a fact that it is nonsense, because at some point we challenged that knowledge by actually putting our hand over a flame. And that's actually what kids do!
As a parent you can tell your kid not to do this or that and they will still do it. I can't recall where I read last week that the most terrible thing about parenting is the realisation that they can only learn through pain... which is probably one of the most efficient feedback loops.
Copilot is no different, it can spit out broken or nonsensical code in response to a prompt but developers do that all the time, especially beginners because that's part of the learning process, but also experts as well. Yet we somehow expect Copilot to spit out perfect code, and then claim "this AI is lousy!", and while it has been trained with a huge body of work it has never been able to challenge it with a feedback loop.
Similarly I'm quite convinced that if I were uploaded everything there is to know about kung fu, I would be utterly unable to actually perform kung fu, nor would I be able to know whether this or that bit that I now know about kung fu is actually correct without trying it.
So, I'm not even sure moving goal posts is actually the real problem but only a symptom, because the whole thing seems to me as being debated over the wrong equivalence class.
Typical LLM AI has been trained for the equivalent of many person-years. How long would it have taken us to read terabytes of information?
> Pretty soon we'll have to accept that we are just biological machines
But what is the code running on those machines?
That's what you are missing. If the mechanics of the machine are a mystery, you will never read the code.
You didn't find the solution, you found your lack of a solution.
No it isn't - his entire argument is that LLM != Humans, just because LLM can do some human like thinhgs. Pointing out the differences isn't moving the goalposts - it's proving the point.
> Pretty soon we'll have to accept that we are just biological machines
Sounds strawmanish, humans != LLM doesn't mean humans == magic.
> and how we respond to our environment is pre-determined
What does this even mean ? Even these models are stochastic.
You seem to be strawmaning that and bringing up almost tautological "once we advance enough towards AGI we will have human level intelligence".
ChatGPT/LLM has a lot of hype - they are fundamentally not human level intelligence.
Some of the biological details are a mystery too. Just recently a 4th brain membrane was discovered:
What I'm getting at is, membranes are cool, I'm not personally very motivated by them.
"This is moving the goal posts"
<insert argument that implies there should be no goal posts>
<insert some uncritical and speculative extrapolation of current AI abilities arbitrarily far into the future>
Honestly, now whenever I see the word "goal post" my eyes glaze over because I already know exactly what's coming next.
Can't find this exact quote anywhere. Who ever said anything like that?
These papers suggest we are just predicting the next word:
https://www.psycholinguistics.com/gerry_altmann/research/pap...
https://www.tandfonline.com/doi/pdf/10.1080/23273798.2020.18...
https://onlinelibrary.wiley.com/doi/10.1111/j.1551-6709.2009...
https://www.earth.com/news/our-brains-are-constantly-working...
Similarly if LLMs can be used to model human intelligence, and predict and manipulate human behaviour, it'll be good enough for corporations to exploit.
If you are correct about LLMs being a generally complete model, then that is a good prediction. But only if you are correct.
What I see you doing here is personifying the model, and drawing conclusions from the personification.
There is more to how we interact with language than prediction of repitition. You didn't predict anything I have said so far! Yet we are both interacting with the language.
We didn't just model LLMs after our brains, either. We pointed them at examples of thought, all neatly organized into the semantic relationships of grammar and story.
Don't ignore the utility of language: it stores behavior, objectivity, and interest.
> This computational organization is at odds with current language algorithms, which are mostly trained to make adjacent and word-level predictions (Fig. 1a)
I feel like you're suggesting because humans != LLMs then humans cannot be doing next word prediction.
I see GP suggesting that because humans!= LLMs that "doing word prediction" is not an exhaustive list of human behavior.
Prediction certainly is one of the things we do with language. That doesn't mean it is the only thing!
It's my contention that most of the behavior people are excited about LLMs exhibiting is really still human behavior that was captured and saved as data into the language itself.
LLMs are not modeling grammar or language: they are modeling language examples. Human examples. Language echoes human thought, so it's natural for a model of that behavior (a model of humans using language) to echo the same behavior (human thought).
Let's not forget, as exciting as it may be, that an echo is not an emulation.
This is not how science works.