I put a grammar error into the last part of the last sentence purposefully, after I wrote it so you can read it and notice the error explicitly. (Whereas I would expect you to still figure out what I meant.) If it were just pattern matching that grammatical error wouldn't bother you at all even if you notice it, you would just go with the most likely fuzzy interpretation.
Because in English word order is important (in fact, in most languages)
It would be fun to learn a completely independent word order language (I suspect Latin is like that) but I suspect this is one of the reasons this aspect died in romance languages, it's harder to understand and relying on word order makes it easier
We kinda ignore mistakes when they don't clash with our heuristics (nobody reads every word, especially stop words)
You can see how people make mistakes by writing the way the words sound rather than constructing the phrase (and this happens accidentally as well, not because the person doesn't know the right words)
Your confidence threshold of the meaning of the sentence was reached, so you didn't wait for the rest of the results to come back before moving onto the next sentence.
Had you needed to understand the sentence sequentially, you would be less likely to miss the grammatical error because you wouldn't be able to skip the word.
I actually had to read the sentence about 3 times, and force myself to read it slowly word by word, before I spotted it.
We often don't read one word at a time.
I agree with you that "massively parallel" is probably drastically overstating it, though.
Think of a single-core processor executing instructions one at a time. Now think of ten billion neurons in your brain out of which a large proportion is active at any time. That is definitely massively parallel.
This is like pointing to the billions of transistors in a typical modern CPU and claiming it is massively parallel despite being single core.
For many things the brain clearly can do multiple things in parallel (I can write one thing while talking about something else, for example). But it most certainly can't usually carry out many high level tasks at the same time.
When it comes to things like recognising words/symbols the question is how high level that actually is, and whether or not the brain has capacity for recognising more than one at the same time, and if so how many.
For fun, I once had two friends speak to me simultaneously, one in each ear. We swapped places and did that to each other a few times. Result: nope, our language understanding unit is not parallel. You can either take in a wall of sound, OR focus to a string of words and understand its meaning. It might feel like you can almost do both at the same time and pick up two different threads of conversation at the same time, but it's always just beyond your reach. At best you can swap between threads quickly, if the conversations permit it (so forget it both/all conversations are rapid-fire and heavy with nuance etc).
I also note that our language generation units are not parallel. We can utter one syllable at a time. So maybe there's not much point in understanding multiple syllables at a time?
Several years ago Kathy Sierra wrote a great post about how app developers should try to educate users from the first time they use the product and "raise the resolution" on how they think about the problem the app is designed to solve.
http://headrush.typepad.com/creating_passionate_users/2005/1...
Well, OK, that's one big question.
However, I can process music and language at the same time. I think I can process multiple pieces of music at the same time, for example I can follow a complex piece of classical music or jazz with multiple instruments playing at various signatures and so on (full disclosure: I'm a drummer). I can't quite do that with speech though. So maybe music and language are just different things and they're processed in different ways?
Your experiment seems to rather point to the intuition that we don't have a plurality of these units which can operate independently, not that a single unit doesn't have a parallel architecture.
The parallel talk is being superimposed into effectively a single input going into the same unit. That unit doesn't have the processing layer to unravel this superimposed input into two streams of speech.
I'd say because the probabilistic structure of language is such that it must be processed sequentially to be decoded correctly. The interpretation of what comes next depends on what came before. Information-theoretically, this happens to be an efficient and compact way of encoding information in the auditory medium (where input unfolds over time, rather than being presented all at once like in images).
Whether processing is "massively parallel" depends on what level of analysis you're assuming. At some level, processing in the brain is obviously massively, since that's simply how the brain's laid out.
I don't see how sequential processing bears any implication to whether or not language is or can be learned through statistical pattern matching, but I'd love to hear your angle on this.
The opportunity for massive-parallelness is in disambiguating a huge number of possible interpretations. Before you can construct the possible candidates, you first need to have processed enough of the sentence. But since you can process new sensory input at the same time as disambiguating previous input, you can get the garden path effect. What I find interesting about the garden path effect though is that it's so artificial, so the way language is used normally seems to purposely avoid ambiguities that will trip up listeners/readers.
Consider this text: "Well the whole point of it was... I mean, when I was in school, we often-- sometimes we would go hunting after school, and I mean, you know -- look, nothing like this ever-- we could never have imagined it."
That's 5 or 6 sentence fragments, one after another. If you processed it in a naively serial way, you just end up with one big ungrammatical sentence. I'm only able to clarify it with punctuation (mdashes and commas) because I could parse it correctly.
A parallel system would benefit not from parsing such a string, but creating a set of candidate interpretations, (all the words together, this is a series of three phrases, this is a series of five phrases, etc) and determining the most meaningful match.
I don't read one word at a time, but scan larger chunks. For instance, I processed "if it's so massively parallel" simultaneously as one "picture".
Do you really read one word at a time? Try making a sheet of paper with a cutout that is approximately word-sized, and move that peephole over the text. Then you're reading one word at a time.
If that makes no difference to your reading efficiency, you could have some cognitive disability affecting reading.
I wonder if the serial limitation is purely historical, or if our intelligence system is somehow fundamentally unable to process massively parallel transmission, even if such has been available during our evolution?
Certainly, we recall associations in a seemingly massively parallel fashion, but our conscious thought is apparently serial. We think this, then that... wait, isn't this really that? There may be non-deterministic threads going in the background, but they are more about association or "fit" than working something out.
We also model things narratively, a serial representation which is also a nice match for events over time - also serial.
BTW: didn't spot the error at first. But I've seen typos in books on second or later readings, which (evidently) also eluded the author and copy editor.
How do we really hear? How do we hear speech? Though the audio is constantly in oscillating motion, tracing an amplitude graph in time, we perceive audio as snapshots: images are formed in our mind which persist. We chunk the audio into frames. Speech understanding probably doesn't take place until a whole frame of audio is assembled and delimited into phonemes. That speech frame appears in the mind as a unit which is randomly accessed; we do not imagine we have a "read write head" which can only retrieve one phoneme at a time, and only in one direction.