On Chomsky and the Two Cultures of Statistical Learning (2011)
norvig.com
norvig.com
C: Your statistical models are woefully inadequate at describing language.
N: That inadequacy is related particularly to Markov models and Ngram models. More sophisticated statistical models will be adequate.
C: Then why haven't you built the more sophisticated models? Why are you still using Markov models and Ngrams?
N: Those work well enough for engineering applications.
The attitude of "it works well enough for engineering" is what Chomsky is actually criticizing. And that criticism is entirely valid: an empirical scientist would never claim that a theory is true because it can be used in engineering.
It's funny to me that Norvig holds up the PCFG as an example of a new and improved statistical model of language. The PCFG is actually terrible in many ways, the most obvious of which is that it doesn't take into account the Theta Criterion [1], one of the most fundamental phenomena of language. An example of this rule is that a noun phrase can only have one determiner. This restriction is so strong that it will never be violated in any kind of professionally composed text. But it is very awkward to try to encode this rule in a PCFG (you essentially have to split the NP symbol into DetNP vs UndetNP). I wrote a blog post describing the problems of the PCFG formalism:
https://ozoraresearch.wordpress.com/2017/03/17/chuckling-a-b...
Some video clips of this interview by Yarden Katz: http://yarden.github.io/pages/chomsky/
> A probabilistic model specifies a probability distribution over possible values of random variables, e.g., P(x, y), rather than a strict deterministic relationship, e.g., y = f(x).
For example, by these definitions, we could use a linear regression as a statistical model that is not a probabilistic model; we could make a bayes network (using known distributions) as a probabilistic model that is not a statistical model; and we could make a Hidden Markov Model trained on sample data that would be both a statistical and probabilistic model.
> y_t = x_t + e_t
where e_t is some error term. Would this be a statistical or probabilistic model?
We do know that the best predictor is just the conditional expectation (in a linear setting bla bla)
> y_pred = E[x_t + e_t | .. ] = x_t
Or is this what you mean with "model"? The predictor? Sorry for being a bit confused.
The regression E[Y|X] is basically the mean of a gaussian distribution of Y given X with sigma set to the error term. The whole gaussian distribution part is the probabilistic model. But to estimate wich parameters make up this model in a particular application (together with how to check if it is valid or not) is the statistical modeling part.