I've been actively learning Portuguese and Russian lately, it's impressive how much faster I can pick up Portuguese vocabulary vs. Russian. And that's even for words that don't have an obvious cognate in languages I already know. The structure of the words, the various building blocks are just so much more familiar in Portuguese. A word like "atrever" (to dare) doesn't have any obvious cognate in languages that I know but it just "looks right" it a way that, say, "atverer" or "aterver" wouldn't. Those last two words sound distinctly un-Portuguese (I might say, un-Romance) to me. That makes it a lot easier to remember the spelling.
Eventually as I grow my Russian vocabulary I start making similar connections. Волноваться is pretty tricky to memorize on its own but it becomes easier when you know that ~ся is the reflective, ~ть is the common verb ending, волна means wave and ~овать is a very common building block for Russian verbs.
Steven Pinker's book Words and Rules is a great layman-oriented read if this sort of thing interests you.
So, they're like "constants" vs calling a function to calculate a value.
I think a decent real example of that is fiancé/fiancée, those are french borrowings and have, at least originally, kept the French grammatical gender inflection. However nowadays I often see people using either spelling in a gender-neutral way since most people don't bother to learn French grammar for this one word.
That still wouldn't explain the why of having it like "went" vs "goed".
Sure, it's easy to memorize because we use it all the time, but why have it to memorize it in the first place, versus something like "goed".
So, this theory (I tried to convey above) said that it being irregular placed ensured we don't slow down try to derive it from regular rules, but instead have fast access to a memorized form.
Couldn't we just memorize "goed"? If it's just "frequency of use" that mattered, "went" and "goed" would work just as well.
But the extra idea is that "goed", being regular, would be too easy for us to confuse with thousand of other regular verbs, and not use our "fast recall" mechanism, regardless of that verb being needed all the time.
Not sure if correct - read it years ago. This seems to be related to that:
https://en.wikipedia.org/wiki/Regular_and_irregular_verbs#Li...
Natural languages are not designed, they evolve. The irregular nature of the conjugation of "to go" might just be a remnant of some archaic form and nothing else, in the same way some argue that the plural of "octopus" is "octopi" or "octopodes. Does it serve any linguistic purpose? I don't see how, but it won't stop some people for saying it. Look at the use of the subjunctive in "if I were you", which is one of the only occurrences of the subjunctive outside of set-phrases in day-to-day modern English. Is it really necessary? If one was to say "if I was you" would it lose some additional nuance or meaning in practice?
I often see people trying to rationalize some language features (such as arbitrary genders of nouns in many languages) as error correction or some "optimization" but I'm generally unconvinced. Maybe "I went" is just the linguistic equivalent of a platypus, some weird byproduct of a very long evolution with no other intrinsic purpose in the grand scheme of things.
But that's part of my point. I don't say irregular verbs were designed to be effective that way, but that they were evolved to be effective.
Hence the link with "language acquisition" being involved -- when regular and irregular verbs developed that wasn't a known theory some "language designer" could consciously follow. Just something innate that could develop because of a evolutionary advantage.
In fact, if someone merely designed, they'd probably go for all regular, rather than regular + irregular, as it's "cleaner".
>I often see people trying to rationalize some language features (such as arbitrary genders of nouns in many languages) as error correction or some "optimization" but I'm generally unconvinced.
Tons of language features are indeed optimizations for different things. Cold climates for example have languages with less vowels (keeping your mouth closed more).
It wouldn't have anything distinctive for people to latch on to, so they would be constantly trying to derive it from the general rules for regular verbs.
E.g.
(a) all verbs regular -> instinctively go to (slower) rule derivation instead of memorization of all verbs, even for the most common ones.
(b) most frequent verbs being irregular -> instinctively retrieve them from the (faster) "lookup table" of memorizations, and bypass the rule based derivation for them.
I.e having the clear distinction of irregularity makes it faster to go directly to that kind of "constant" memory.
That said, this is not my theory, read it years ago in a cognitive/linguistic pop science article. This seems to say more or less the same thing:
https://en.wikipedia.org/wiki/Regular_and_irregular_verbs#Li...
In studies of first language acquisition (where the aim is to establish how the human brain processes its native language), one debate among 20th-century linguists revolved around whether small children learn all verb forms as separate pieces of vocabulary or whether they deduce forms by the application of rules. Since a child can hear a regular verb for the first time and immediately reuse it correctly in a different conjugated form which he or she has never heard, it is clear that the brain does work with rules; but irregular verbs must be processed differently.
~ов is an iterative suffix - like ~le or ~er in gamble and chatter. It's useful to know, because you can rationalise why it is always dropped in the present tense - you can't be iterative at the moment, unless you're an Englishman)
I didn't know that, but you'll grant me that it's not a very useful cognate (either in spelling or in meaning).
>~ов is an iterative suffix - like ~le or ~er in gamble and chatter. It's useful to know, because you can rationalise why it is always dropped in the present tense - you can't be iterative at the moment, unless you're an Englishman)
Very interesting, thanks.
Converting a large word list to a small number of bits has been a computer science hobby for a long time. Here's a pretty good search result to start working through for more details: https://duckduckgo.com/?q=building+a+small+spell+checker+suf... It was especially important to write small and fast spell checkers in the 1980s and early 1990s, when you couldn't expect to have enough RAM sitting around to simply load up a naively-encoded list of words, and the act of spell checking a few thousand words could take noticeable time.
So in an encoding scheme chosen to represent English compactly, I'm not too surprised that you can get things down quite small.
However, the question is, what relevance does that encoding scheme have to the human brain? Having just scanned through the paper, the answer is "probably not that much", which the researchers are well aware of. They explicitly present this as a lower-bound, which is a reasonable thing to do. It is obvious that the brain does not simply store 1.5MB of data in the way a computer does, in many ways.
To be honest, this amounts to an exercise in recreational mathematics more than anything else. There's nothing wrong with that, and that's not a criticism. My point is that I'm not sure it's worth trying to read the paper as anything else.
Now a few questions: Can you hear a word you've never written (in a language that you're familiar with) and intuitively spell it right the first time? Can you read a word that you don't immediately understand and figure out its meaning from the context in which it is used? Can you accurately complete half of a sentence?
That a lot of people can do these things suggests to me that we all sit on something superficially similar to an efficient lossy compressor in our brain.
Considering that, 40k words to 400kbits is not too surprising.