The choice of the use of the sigmoid function does not stem from applying MLE to the model: instead it is an priori and arbitrary modeling decision that has been done even before starting to think about the estimation of the parameters of the model.
The intuitive justification with respect to the choice of this link function given in this blog post seems quite standard to me. This text-book chapter gives a similar intuition: http://www.stat.cmu.edu/~cshalizi/uADA/12/lectures/ch12.pdf (section 12.2).
Also in practice we don't use pure MLE but l2-penalized (or l1-penalized) MLE to fit a logistic regression model as otherwise it might be very sensitive to noisy data if the number of features very large (compared to the number of training samples).
Thanks for pointing out the relation to the Bernoulli distribution.
The "logistic regression" has been invented / reinvented independently with different derivations 10 times or so.
Which is to say that the starting point is that of someone figuring out what is going on behind the scenes of a working software package rather than building up from first principles.
Curious if there is a link to a better explanation from a similar starting point.
Maybe a better analogy might be cryptography where practical implications of entropy pool design are relatively esoteric despite the vastly less accessible mathematical nature of reliable one way encryption algorithms.