The idea behind binary regression ( Y is 0 or 1) is that you use a latent variable Y* = beta X + epsilon.
X is the matrix of indenpent variables, beta is the vector of coefficients and epsilon is an error term that sums the rest of what X can't explain.
Y thus becomes 1 if Y* >0 and 0 otherwise.
Seeing how Y is binary, we can model it using a Bernoulli distribution with a success probability P(Y=1) = P(Y* >0) = 1 - P(Y* <=0) = 1 - P(epsilon <= -beta X) = 1 - CDF_epsilon(-beta X)
Technically you can use any function that maps R to [0, 1] as a CDF. If its density is symmetrical then you can directly write the above probability as CDF(beta X). The two usual choices are either the normal CDF which gives the Probit model or the Logistic function (Sigmoid) which gives the Logit model. With the CDF known you can calculate the likelihood and use it to estimate the coefficients.
People prefer the Logit model because the coefficients of the model are interpretable in terms of log-odds and all and the fucntion has some nice numerical properties.
That's all there is to it really.