Critical Behavior from Deep Dynamics: A Hidden Dimension in Natural Language
arxiv.org
arxiv.org
BTW, I went to the North American Association of Computational Linguistics conference in April and it seemed like half the papers used LSTM.
Edit: the NAACL 2016 papers are here: http://aclweb.org/anthology/N/N16/
As someone outside of this field, it seems to me that this kind of result should have been very obviously foreseeable, hindsight bias and all that of course - but I would never have considered Markov processes to be an adequate predictability model for natural language. Though obviously the formalized results are important.
Could someone with more knowledge comment on what the current working assumptions were prior to this paper and what the consequences would be?
Would you consider LSTM an adequate model?
That doesn't mean they're not useful in very narrow domains. But language is pretty much the definition of the ultimate wide domain, and trying to cover it with statistical correlations makes as much sense as word counting Shakespeare to try to generate some new plays.
The relevant example from the paper:
I: Jane went to the hallway.
I: Mary walked to the bathroom.
I: Sandra went to the garden.
I: Daniel went back to the garden.
I: Sandra took the milk there.
Q: Where is the milk?
A: garden
Obviously just a toy task, but as you said, progress is rapid!What exactly those models are modelling?
From my non-professional perspective the above seems that it should have been very obvious - (and also that correlations between variables for natural languages would be better explained by multi-dimensional structure). That is if you told me that this were proved / formally supported as it is in this paper, my reaction would be a "that sounds like reasonable approximation" not - "that result sounds very surprising I must read the paper".
Isn't it the case that for very short distances (several elements), power decay and exponential decay are (or can be made, with proper constants) more similar? Thus, if predictive models were originally studied only for very short sequences in the past (limited computational resources!), it seems to make sense that this is a mistake that anyone could have made more easily back then.
> mark_l_watson
Hah, I'll take your word for it, then. :) Are there any recent comprehensive monographs you'd recommend for state-of-the-art NPL, for someone who has yet to enter the field?
* Foundations of Statistical Natural Language Processing by Manning and Schütze
* Speech and Language Processing by Jurafsky and Martin (which is being revised for a third edition, which you can look at: https://web.stanford.edu/~jurafsky/slp3/ )
Beyond that, you're basically stuck reading the research literature. On the up-side, most of that literature is freely available from the ACL anthology at http://aclweb.org/anthology/
Ah yes, the one I've heard about but still have to take. :) Well, I guess I should give Coursesa a chance. (Somehow I'm not fond of their "timelined" format, it seems redundant if you're communicating with a machine. I hope the future of online learning will avoid it like the plague.)
I did not read it very closely though.
> We can formalize the above considerations by giving rules for a toy language L over an alphabet A. In the parlance of theoretical linguistics, our language is generated by a stochastic or probabilistic context-free grammar (PCFG) [41–44]. We will discuss the relationship between our model and a generic PCFG in Section C.
Mmmh
> [...] We show that in many data sequences — from texts in different languages to melodies and genomes
Hum, ehrm
> [...] natural languages are poorly approximated by Markov processes.
Alright, alright
> [...] This model class captures the essence of probabilistic context-free grammars
Ok, ok
> [...] and cosmological inflation
Wat.
Out of nowhere, Creation of the Univerve.
-------------
I'm always baffled by the ability to draw parallels. Did a colleague take at peek at the screen and said, hey I have the same equations?
"...and that, my liege, is how we know the Earth to be banana-shaped."
"This new learning amazes me, Sir Bedevere. Explain again how sheep's bladders may be employed to prevent earthquakes."
Max is deeply invested in modeling, analysis and prediction software, and I suspect did the bulk of the work in the paper.
Henry Lin is a student who is focused on astrophysics. He gave an interesting TED talk (http://www.ted.com/speakers/henry_lin) a few years back about studying distant galaxy clusters.
Henry is energetic and almost viscerally inspired by the beauty of science and mathematics, such a wonderful quality! His voice is definitely in the prose of the paper.
This would an interesting thing to try - a computer system that would scan all the papers for math and find parallels. I think we already have something like term indexing for deductive systems?
That might be a bit of a narrower domain, but it seems to work out pretty well!
Corollary: No probabilistic regular grammar
exhibits criticality.
In the next section, we will show that this statement is
not true for context-free grammars (CFGs).
That is, there exists CFGs that exhibit criticality. Programming languages are often parsed by CFGs, so it's likely that some programming languages exhibit the same criticality structure as natural languages.For reasons like this, I don't trust their empirical results at all.
-----
you: tell me something
me: Don't speak for me first time I asked it
you: tell me about your mother
me: just like my mother is on that AK47 diet
you: Tell me about artificial intelligence
me: Okay , maybe not intelligence capabilities, etc
you: ask me something
me: Points bow No one would ask this haha
Is that a joke? It's 2016 and you think you need to explain that a Markov process is a poor approximation of natural language? This has been obvious for computational linguists, and anyone working in the field, from day one.
The Bach data consists of 5727 notes from Partita No. 2 [11],
with all notes mapped into a 12-symbol alphabet
consisting of the 12 half-tones {C, C#, D, D#, E, F, F#, G, G#, A, A#, B}
with all timing, volume and octave information
discarded.I was good until the last part.