Maximum likelihood estimation and loss functions
rish-01.github.io
rish-01.github.io
>Let pmodel(x;θ) be a parametric family of distributions over a space of parameters ...
and it's straight to the grad-level textbook stuff that breezily assumes familiarity with advanced mathematical notation.
One of the reasons I loved Andrew Ng's machine learning course so much is that it eased you into understanding the notation, terminology, and signposted things like "hey this is really important" vs. "hey this is just a weird notational quirk that mathematicians have, don't worry about it too much."
I used to get by on, “it’s the parameters that make the data most likely”, like it says on the name. I think that’s what you are after.
Then I took a stats class, and I know to say, “the MLE is minimum variance within the class of asymptotically unbiased estimators” … that is, “efficient” and “asymptotically consistent“ in the jargon. (Subject to caveats.)
Then I took a Bayesian stats class and learned to say, “it’s minimum risk under a (improper) uniform prior.”
I also recall there is a general result showing that any estimator which makes the score function zero has good properties with respect to average loss. So zero’ing the score by maximizing likelihood is a good strategy. (If someone could remind me of specifics, that would be great.)
But perhaps Gauss had it right when he exploited the (known, but yet un-named) central limit theorem and used how easy it is to maximize the quadratic that sits atop its “e”. (https://arxiv.org/pdf/0804.2996, page 3, top). It’s so easy we had to find a justification for using it?
Note that you don't need to go deep into measure-theoretic probability or any of that stuff that requires more advanced prior education in math.
Perhaps a first course at grad level, but my engineering bachelors covered MLEs but we didn’t learn/use any of those formal language things. I think the core mathematics (and likely other pure science) cohorts were the only people who learnt it.
It took six months for Math 100 (Maths for wanna be mathematicians) to "catch up" with the applications being spat out in Math 101 (Maths for people that practically use math for applications) but by the time the foundations were laid almost all the applied math in the Engineering coursework results just became "an exercise for the reader" to derive without need for rote memorisation.
My courses definitely used different notation for the same semantics.
At first glance likelihood functions might look the same, but you have to think of them as functions of the parameters; it's the data that's fixed now (it's whatever you observed in your experiment). Once that's clear, the calculus starts to makes sense -- using the derivative of the likelihood function w.r.t. the parameters to find points in parameter space that are local maxima (or directions that are uphill in parameter space etc).
So given a model with unknown parameters, the data set you observe gives rise to a particular likelihood function, in other words the data set gives rise to a surface over your parameter space that you can explore for maxima. Regions of parameter space where your model gives a high probability to your observed data are considered to be regions of parameter space that your data suggests might describe how reality actually is. Of course, that's not taking into account your prior beliefs about which regions of parameter space are plausible, or whether the model was a good choice in the first place, or whether you've got enough data, etc.
The reason it's not a probability or probability density is that it's not defined to be one (in fact its definition involves a potentially different probability density for each point in parameter space).
But I think I know what you're saying -- people need to understand that it's not a probability density in order to avoid making naive probabilistic statements about parameter estimates or confidence regions when their calculations haven't used a prior over the parameters.
I know it's not statistically correct but I think it helped a lot in my understanding of other methods....
Let's say you have some guarantee that I'm using the same die each time and that each of the rolls are independent. We play the game ten times and 1 is rolled the first 9 out of 10 times, with a 5 being rolled on the 10th throw. Now, you know that there's a common loaded die that can be purchased that has a weight to skew the probabilities and you further know that the loaded die rolls a 1 80% of the time and the remaining 20% spread evenly to the other values (so 4% for every other value).
Given a choice between the loaded die and the fair die, which is more likely?
The first model, call it $\theta_0$ is the fair die. The second model, with the unfair die, call it $\theta_1$.
The probability of the first model ($\theta_0$) is:
$p( 1,1,1,1,1,1,1,1,1, 5 ; \theta_0) = \frac{1}{6}^9 \cdot \frac{1}{6}$
(approximately .00000008269085843959)
The probability of the second model ($\theta_1$) is:
$p( 1,1,1,1,1,1,1,1,1, 5 ; \theta_1) = (0.8)^9 \cdot 0.2 \cdot \frac{1}{5}$
(or .00536870912000000000)
So we write a computer program to iterate through all the "models" to see which is the more likely. In this case, the iteration goes through two models.
The models can be Gaussians, with model parameters the mean and variance, say, or some other distribution with other parameters to choose from.
For some conditions on models and their parameterization, we might even be able to use more intricate methods that use calculus, gradient descent, etc. to find the MLE.
The MLE formalism is trying to say "given the observation, which parameters fit the best". It gets more complicated because we have to talk about which distributions we're allowing (which "model") and how we parameterize them. In the above, the models are simple, just assigning different probabilities to each of the outcomes of the die rolls and we only have a choice of two parameterizations.
Technically, these functions need to satisfy a bunch of properties, but those properties matter mostly for people doing the business of building and comparing models. If you just have a model someone already made for you, then "the higher the better" is good enough.
It's also the case that these models have "parameters". As a simple example, the model of a coin flip takes in "heads" or "tails" and returns a number. The higher that number, the more probable it claims that outcome to be. When we construct that model, we also choose the "fairness" parameter, usually setting it so that both heads and tails are equally likely.
So really, it's a function both of the data and of its parameters.
Now, "maximum likelihood estimation" (MLE) is just the method where you fix the data inputs to the model to whatever your training data is and then find the parameter inputs that maximize its output. This kind of inverts the normal mechanism where you pick the parameters and then see how probable the data was.
Presumptively, whatever parameterization of your model makes the data the most likely is the parameterization that best represents your data. That doesn't have to be true, and often is only approximately true, but that presumption is exactly what makes MLE popular.
Finally, it's worth describing the origin of the name. When we look at our model after fixing the data inputs and consider it a function of its parameters instead we call that function a "likelihood". This is just another name for "probability" except it's used to emphasize that likelihoods don't meet all the technical properties I skipped up above. So "maximum likelihood estimation" is just the process of estimating the parameters of your model by maximizing the likelihood.
This intuition really helped me understand CE loss.
https://stats.stackexchange.com/questions/357963/what-is-the...
Please correct me if I'm wrong.
To solve the equation, we have to make assumptions of the poor correction factor. These assumptions about the error generally have some 'mathematically nice' qualities. For example it's not predictable or has a trend relating to any other factors. An concrete example is having a mean of zero. If it had a non-zero mean, it should be accounted in the constant factor of the model.
All these mathematically nice assumptions can be summed up be calling the 'poor model correction' factor as random.
- Say I have data which is perfectly sinusoidal, with an dc bias. I can fit a line a to this data, which will approximate the bias (or be exactly the bias if the data is over an integer number of cycles).
- I want to fit a plane to a curved surface
- I want to fit a low order transfer function to a high order system.
- I want to model a system with friction as a system with no friction.
Fitting parameters in all of these situations will result in a non-zero residual. But assuming that is due to randomness is not useful.