A sane introduction to maximum likelihood estimation and maximum a posteriori
blog.christianperone.com
blog.christianperone.com
OT on technical blogs: Experts often are unable to put themselves in the shoes of someone with no experience, which really harms the pedagogy. When one practices a technical topic for a long time, concepts that were once foreign and difficult become instinctual. This makes it very hard to understand in what ways a beginner could be tripped up. It takes a large amount of thought to avoid this problem, which I think is why much introductory material - blog posts, books, etc., is really sub-par.
If anyone is interested in learning more, this phenomenon is typically called "the curse of knowledge".
Sorry if this is obvious but I have been doing a lot of reading on this and have come across this step a few times before...but am just missing some part of every explanation.
One note though: I think on equation 25 you are missing a log on the left hand side.
Typically we write the likelihood function as
L = Π P(y | θ)
If you didn't have identically-distributed observations, the functional form of P would be different for each observation.And if you didn't have independent observations, then you're basically screwed in the general case. That expression for L is basically the definition of probabilistic independence: a finite set of random variables is mutually independent if and only if their joint probability function is equal to the product of the individual variables' probability functions.
If you have dependence between observations, you lose the ability to write L in that nice form. This is a non-negotiable consequence of basic probability theory.
The only way to do MAP estimation without iid observations is to know the joint distribution of your entire dataset, and be able to maximize that distribution with respect to θ given an arbitrary data set. This is possible but it's not quite the same thing as dumping your data into a GLM.
In fact, not approaching the more general case is liable to confuse learners as they may think that independence assumption is somehow baked into MAP/MLE, which it is not.