A Simple Structure Unites All Human Languages
nautil.us
nautil.us
Also: in preparation to diving in Kafka's books I learned about a peculiar feature of his style:
> Kafka often made extensive use of a characteristic particular to the German language which permits long sentences that sometimes can span an entire page. Kafka's sentences then deliver an unexpected impact just before the full stop—this being the finalizing meaning and focus. This is due to the construction of subordinate clauses in German which require that the verb be positioned at the end of the sentence.
>> “Als Gregor Samsa eines Morgens aus unruhigen Träumen erwachte, fand er sich in seinem Bett zu einem ungeheuren Ungeziefer verwandelt.” (original)
> “As Gregor Samsa one morning from restless dreams awoke, found he himself in his bed into an enormous vermin transformed.”
There's a neat picture illustrating the difference in the order of the parse tree: https://en.m.wikipedia.org/wiki/Franz_Kafka_bibliography#Eng...
From the Wikipedia article:
> German also lacks an informal language register
Can someone provide more insight into what this refers to? There are definitely less formal or technical sounding word variants in German, and of course duzen/siezen to add another level of formality, so I'm not sure what this could refer to.
I'd assume it's about the dialects, which are very different from the "official" language, so much that Jerome K. Jerome had a joke in one of his books, more than century ago:
"Germany being separated so many centuries into a dozen principalities, is unfortunate in possessing a variety of dialects. Germans from Posen wishful to converse with men of Wurtemburg, have to talk as often as not in French or English; and young ladies who have received an expensive education in Westphalia surprise and disappoint their parents by being unable to understand a word said to them in Mechlenberg."
"Modern times" and technologies suppress the dialects and with each generation the portion of the local-specific dialects is being lost.
So, afaiu an ‘informal register’ would be something like brospeak or language spoken at home and among friends, contrasted to that spoken with strangers and at work. But I don't know what the situation is in German. With English and Russian, every generation and each subculture invents their own slang just to differentiate themselves―can't imagine how any country would avoid developing informal language, considering the existence of Oktoberfest.
Could you (low register to someone of equal or lower status) vs Could you (medium register to someone of higher status) vs Might you (highest register to someone of higher status). In English I had to change to a different verb, non grammatically. In German these are all the same verb with a different rule applied.
You know, this actually makes sense. I think that after a couple of hundred pages, this would just seem like an easy-to-read, natural alternative word ordering.
Then, if the author started substituting German words here and there, starting with the obvious ones, then ones English words are derived from, and so on... before you know it, you'd be reading German!
> Word order is, of course, is far more complex than I’ve shown here. There are languages with very free word order, and even within languages there are many intriguing complexities. However, this idea, that Merge can both combine bits of language, and reuse them, gives us a unified understanding of how the grammar of human languages works.
But none of the examples in this simplified account even gestured at noun case, or at the prospect of expressing subject (or agent) with verb conjugation, or at feature agreement.
Is there a straightforward way to understand why Chomsky thinks that this approach addresses those phenomena?
Also is 'Merge' a thing? Or has the author just blessed his concept into proper noun-hood? It's not really explained.
"'Syntactic Hierarchical Structure All the Way Down' entails that elements within syntax and within morphology enter into the same types of constituent structures (such as can be diagrammed through binary branching trees). DM is piece-based in the sense that the elements of both syntax and of morphology are understood as discrete instead of as (the results of) morphophonological processes."[1]
So a rather simplified way to think about it is that each word is a little mini-tree and Merge operations create a branching structure between its constituent parts. Each of those words is then part of the larger sentence-level tree. The important this is that Merge is acting on the units at both levels.
Similarly, Merge can act on structures that are the output of previous Merges, allowing you to have a verb, in a particular conjugation (it's sub-word structure), that selects for a particular type of supra-word structure (a tree) that's headed by a noun with some set of features, eg a particular case.
Another point that is kind of alluded to in the article is that you can create movement from one part of a tree to another with Merge. In previous theories of syntax, there always needed to be both something like Merge and a special "move" operation. But Merge simplifies things quite a bit in that regard.
[0]https://en.wikipedia.org/wiki/Distributed_morphology
[1] https://www.ling.upenn.edu/~rnoyer/dm/#how%20DM%20is%20diffe...
Can you explain a little more about how something like an agreement rule would be analyzed in this framework?
I suppose I didn't quite understand your "that selects for a particular type of super-word structure (a tree) that's headed by a noun with some set of features". Is this sort of akin to a type system in programming? Like the verb is only willing to bind with a subject noun phrase whose head has a particular feature?
I think using a language which doesn't care much about word order will be illustrative here, so let's use Latin:
puella vidit canem
The girl sees the dog
We can break this down into:
puell - a vid - et can - em
girl - NOM.S.FEM sees - S.3 dog - ACC.S.FEM
So 'puella' is the tuple of features [+NOM, +S, +FEM], 'videt' is [+PRES, +ACTIVE, +S, +3], etc.
Here, we want to do a Merge with 'puella' and 'videt': we say that 'puella' selects for the features +NOM, +3, +S (nominative, third person, and singular) in its verb, but doesn't care about the others. It can still agree with its verb if the verb is passive or in the past tense. But if a verb is conjugated in a way that violates the features it selects for (eg the verb is conjugated as first person plural), 'puella' won't merge with it.
As you said, a phrase level structure will have the features of its constituent parts bubble up to it. So once we've done the first Merge with 'puella' and 'videt', our structure is now selecting for a noun phrase that has the feature +ACC. Because 'canem' meets this requirement, we can get the final Merge necessary for our finished sentence.
{ { puella, videt }, canem }
Note that this account still works if we change the order of the sentence to any configuration, we just need to reorder the merges.
mala mala mala sunt bona
Soli soli soli
This model really can account for quite complex language data though. For example, check out this account of auxiliary verbs in Basque: https://www.academia.edu/3112898/A_Distributed_Morphology_An...
Speaking to your examples: "mala mala mala sunt bona" isn't particularly difficult to analyze this way, you just need to realize that the "mala"s are different words (kind of like the famous English "Buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo" sentence). If I remember the proverb correctly, it means "apples (mala) are good (sunt bona) for a painful jaw (mala mala)".
You need an analysis that allows adjectives to Merge with nouns iff they match case, gender and number, so that allows us to create a noun phrase "mala mala" in the instrumental ablative. Then you need a way to have the case, gender and number of subject bubble to the top of the phrase it will make with an auxiliary so that the adjective after the auxiliary is feature restricted to that case, gender and number. Once the elements of the auxiliary verb phrase have Merged, you get:
{{ mala, sunt }, bona }
Finally you have a rule that allows auxiliary verb phrases to Merge with noun phrases headed by an ablative. If you want the first "mala" to be the subject, then re-Merge it with the whole sentence so far, which in effect moves it to the top of the tree, leaving a trace in its original position.
I'm not sure what the second example means. My best guess is that it's the dative singular of 'sol', a matching masculine dative singular of 'solus' and a genitive singular of 'solum', so something like "for the only sun of the land". If that's correct, you need our previously used rule for Merging adjectives iff they match the noun in case, gender and number. Then you can add an additional rule that genitive nouns can be Merged with noun phrases (without any feature selection needing to take place) to form a new noun phrase.
Hopefully that shows that Merge and feature selection as mechanisms can be used outside of toy models, to actually account for real data.
Anyway, yes, the Merge and feature can work just fine outside of "toy models" the note was about about they soon becoming complex.
I suppose that you could, for example, account for the different conjugations and declensions by saying that they are also features of noun and verb stems that have to agree with endings that want to bind with them, right? Like "vid-" and "-et" is not just "sees - S.3" but also something like "see [+2conj]" and "S.3 [+2conj]" allowing them to bind with each other, where "-at" might be "S.3 [+1conj]" so it could bind with "am-" being "love [+1conj]", while "-et" doesn't bind with "am-" (except when interpreted as a different lexical item that adds [+subjunctive] to a [+1conj] stem?).
My next question is whether there are tools to facilitate writing parsers with this framework because it makes me want to write a Latin parser and see how well it does (and maybe how many formal syntactic ambiguities exist in Latin texts that we might not even notice most of the time).
It's like people look only at their own language and try to pretend that same structure must be universal.
Can you suggest a more example of a linguistic phenomenon that you think this framework can't deal with?
Chomsky isn't concerned with just grammar of written or spoken languages―instead he shows that these grammars map to a universal mental grammar. Different physical languages deal in different ways with constructs of the mental grammar: English expresses the same concepts and relations as inflections, but uses helper words for that. Correspondingly, where you build a grammar tree from words in English, you build it from roots and affixes in inflectional languages. There are languages in which a single word with a bunch of affixes equates to an English sentence.
Regarding “expressing subject (or agent) with verb conjugation”: this should just translate straightforwardly to an implied subject. Even incomplete sentences in whatever language have their implied subjects, verbs and objects―usually deduced from previous speech. Pinker in ‘The Language Instinct’ has a great example of how casual speech is very compact compared to legalese writing where you have to explicitly write out everything in anticipation of ambiguities and adversarial reading.
Chomsky also has a concept of ‘trace’ which is an ‘invisible’ word in a sentence and refers to a previously mentioned word. E.g. in “the spoon that I'm eating soup with,” there's a hidden member: “the spoon that I'm eating soup with <trace>,” and the trace refers to the spoon, helping to build the grammatical tree. I think this concept is hinted at in the article with Gaelic “caught boy caught fish.” Afaiu the ‘trace’ is different from sub-word and implied entities, but it helps to illustrate how the mental grammar isn't the same as written one.
Yes you can express any set as a tree (in fact as many possible trees) but if different possible tree representations of that set are equally as valid - then the tree structure isn't really there, it's just an arbitrary interpretation imposed on the data.
If a sentence A B C can be reordered to any permutation, and some parts can be dropped without changing the meaning - then what lets you decide that the proper representation is this tree:
.
/ \
A .
/ \
B C
and not this: .
/ \
. C
/ \
A B
And if any of these is as valid as any other - why insist that the underlying structure is even a tree?Another hint that these kinds of interpretations are wrong is that most succesfull NLP software typically uses word vectors which ignore the tree-like structure and just treats the data as unordered set :)
To understand a sentence it is more important to know all the words in it, than to know how they are nested. Even for analytical languages like English :)
A good overview: http://langsci-press.org/catalog/book/255
Edit: grammar
That's only true when you think about LSTMs, but stacked CNNs, tree-LSTMs, graph neural net pooling and attention layers can do hierarchical aggregation. Hierarchical representations have been at the centre of many papers. There's even hierarchical reinforcement learning for describing complex actions as composed of simpler actions.
And trees are not good enough to represent language. Graphs would fit better because some leaf nodes in the tree resolve or refer to nodes on other branches (e.g. when you say He referring to the word John present in another place in the same text).
The most helpful thing I think you can do while studying language, other than placing yourself in real world scenarios where you use it, is to just read (or make up) example dialogues. Words by themselves aren't that helpful because they're often out of context. Sentences are better, but even they benefit from being embedded in a larger structure such as a dialogue or paragraph. I guess the point is, the more context, the better. Or, as the article would say, the more merging the better.
to drink wine = "pić wino"
I drink wine = "piję wino" or "ja piję wino" or "piję ja wino" or "wino piję ja" or "wino ja piję" or "ja wino piję"
boy caught fish = "chłopiec złapał rybę" or "rybę złapał chłopiec" or ...
There are many languages in which word order doesn't matter, it's the changes to the words that encode role of the word in the sentence, and you can divide the sentence into many alternative hierarchies (and certainly not all of them are strictly binary - some parts of the sentence are trinary or even more complicated, imposing binary structure on them is artifical and misleading).
I think this theory is very lacking in predictive power, says very little about supposed "universal" language and still isn't really universal as there are all sorts of exceptions.
Natural languages doesn't follow formal grammars strictly, especially not such a simplistic one.
[0] https://golem.ph.utexas.edu/category/2018/02/linguistics_usi... https://arxiv.org/pdf/1809.05923.pdf
http://semantics.uchicago.edu/kennedy/classes/s07/myths/nevi...
https://www.academia.edu/3112859/Evidence_and_argumentation_...
The definition of language is (according to Google), human communication.
For there to be communication, you need a recipient. Which means that mere written words have no meaning without someone reading and interpreting them.
You can analyze syntax and structure all you want, but meaning depends on people's interpretations, which are subjective and depend on multiple other factors.
For instance, if I'm angry, I'll read a text message or an email and interpret it in a completely different way than if I'm calm. The meaning I interpret will also depend on who sent me the message. Neither of these things are captured in the syntax of the messages.
> what is a written message that isn't read?
Imagine an ancient civilization that left written symbols, but the people are long gone, there's no one that knows how to interpret them anymore.
During the time the symbols were not being seen or interpreted by anyone, what would you call them?
But more important than that, is not that the symbols can't mean anything, rather that the meaning will be assigned by the reader when they read it (not just by the syntax of the symbols, which is what the article seemed to imply). And that meaning can be very very different than what the writer intended it to be.
What I'm basically saying is that meaning/interpretation of communication/messages is fluid/dynamic. It depends on the writer, the symbols, the reader and a lot of context. It is not fully contained or captured just by the symbols in which we express it.
Using your comment as an example, your "rereadings are also colored by the contexts shifts".
Question I the depth this analysis of. Paring rule variability or recursive process, or randomized association efficiency? Arbitrary hierarchy inherent world model categories captures actor, act, actee of. New Guinea highlands reference I languages of number large (day before yesterday), exploit possibilities almost all categorical where.
{α, β}
Both α and β can be either some atomic unit, eg a word or a morpheme, or the previous output of Merge (a tree).
This is their first example as rendered as cons:
(cons I (cons drink wine))
As merge:
{ I, { drink, wine } }
As a tree:
Edit: I can't get the tree to look even remotely correct in HN formatting, but you get the idea. It's in the article.
>> [Merge] applies to discrete units of language (words or their parts). It combines these, not sequentially, but hierarchically.
Or the paragraph at the start where human language abilities are contrasted to bonobos' and chimpanzees' and deep learning models' sequential processing; etc.
My knowledge of Lisp is rusty, but as far as I remember it cons is a list operator that joins the head to the tail of a list (like the "|" in Prolog). So it imposes an order - on a sequence. Apologies if I misremember this.
The author is just trying to say that Merge is doing the same thing we're talking about cons doing to S-expressions to a tree and noting that it creates hierarchy. Eg a new Merge says I've created a new top-level node that is a parent for the two inputs. A second Merge says I've made the first input c-command both inputs of my first Merge and created a new top-level node.
If the proposition was Merge as a "universal" operation, then I'd say there is no evidence that our brains implement such an operation, and that it has a very shallow stack, which is domain dependent. That makes such an operation a meaningless abstraction, not suitable for explaining anything about human behavior whatsoever.