Linear regression has a closed form solution of X projected onto Y: \hat{\beta} = (X'X)^{-1} X' Y
It is equivalent to the Maximum Likelihood Estimator (MLE) for linear regression. However, for logistic regression, MLE would estimate different from MLE for the log odds output.
Linear regression on {class_inclusion} = XB gives the linear probability model, which has limited utility. The required transform is covered by another commenter.
The claim is that “evidence” in Bayesian reasoning naturally acts upon the log-odds, mapping prior log-odds to posterior log-odds additively. To see this, calculate from the definition,
Odds(X | E) := Pr(X | E) / Pr(¬X | E)
= Pr(X ∧ E) / Pr(¬X ∧ E)
= (Pr(E | X) × Pr(X)) / (Pr(E | ¬X) × Pr(¬X))
= LR(E, X) × Odds(X),
Where LR is the usual likelihood ratio.So when we take the logarithm of both sides, we find that new evidence adds some quantity—the log of the likelihood ratio of the evidence—to our log of prior probability, in this phrasing of Bayes’ theorem.
I sometimes tell people this in a slightly strange language, I say that if we ran into aliens we might find out that they don't believe the things are absolutely true or false, but instead measure their truth or falsity in decibels.
So another perspective on what logistic regression is trying to do, is that it is trying to assume linear log-likelihood-ratio dependence based on the strength of some independent pieces of evidence. You can weakly justify this in all cases, using calculus and assuming everything has a small impact. You can further justify it strongly for any signal where twice as large of a measured regression variable ultimately implies twice as many independent events at a much lower level happening and independently providing their evidences for the regression outcome. So like, I come from a physics background, I am thinking in this case of photon counts in a photomultiplier tube or so: I know that at a lower level, each photon is contributing equally some small little bit of evidence for something, so when I count all the time up together, this is the appropriate framework to use.
From this point of view linear regression would be using GLM with identity link function, logistic regression uses the logit function as the link function.
In that framework, they are literally the same model with different "settings" - Gaussian vs Bernoulli distribution.
Linear regression is extra-special because it's a special case of several different frameworks and model classes.
I should have written that it's better (in my opinion) to think of logistic regression in the context of GLMs, at least while you're learning.
Edit: yes logistic regression is a special case of regression with a different loss function. But it's not nearly "as special" as linear regression.
Joke aside, the truth is that logistic regression can be understood based on several assumptions.
Above we have the latent variable explanation. Then there is a Bayesian version. There's even a "random utlity" formulation, where one models explicitly choices of an agent with a probabilistic error. That one is good to explain hierarchical logit models and many of the "issues" with logit such as IAA.
GLM on the other hand I don't feel like it adds much except parameterizing the procedure, which ain't even a good thing. Nowadays we appreciate the semi parametric nature of regression a lot, which is why GLM has declined in use.
The point about the probability distribution is reasonable, but I am not sure if it is taken seriously by everyone applying GLM either. And again, if it is not necessary to assume such a distribution, then I would prefer a semi parametric approach, such as in linear regression.
Its not just the underlying model, but the loss function is also different.
Almost - logistic regression assumes that the function is linear in the log odds, i.e. log(p/(1-p)) = Xb + e. The problem is that you can't compute the log-odds, because you don't know p.
log p(y=1 | x; beta) = beta * x - log Z(x; beta)
where
Z(x) = p(y=0 | x; beta) + p(y=1 | x; beta)
Thus, you can think of it as linear regression, but with an additional term log Z(x; beta) in the log likelihood.