NN 201 explanation: mean-square is about finite but continuous errors (residuals) of predicted value vs true value of the output, cross-entropy is for distributions over discrete sets of categories.
NN 501 explanation: the task of the NN and the form of its outputs should be defined in terms of the "shape" or nature of the residuals of its predictions. Mean-square corresponds to predicting means with Gaussian residuals, cross-entropy corresponds to predicting over discrete outputs with multinomial (mutually exclusive) structure. Indeed you can derive any loss function you want by first defining the expected form of the residuals and then deriving the negative log-likelihood of the associated distribution.
"Many authors use the term “cross-entropy” to identify specifically the negative log-likelihood of a Bernoulli or softmax distribution, but that is a misnomer. Any loss consisting of a negative log-likelihood is a cross- entropy between the empirical distribution defined by the training set and the probability distribution defined by model. For example, mean squared error is the cross-entropy between the empirical distribution and a Gaussian model"
https://stats.stackexchange.com/questions/288451/why-is-mean...
Minimizing mean-squared-error loss is equivalent to minimizing KL divergence (and thus cross entropy) under the assumption that your model produces a vector that parameterizes the mean of a multivariate Gaussian distribution which is then used to predict your data. This is the most natural way to set up a model that predicts continuous data.
KL-Divergence is not a metric, it's not symmetric. It's biased toward your reference distribution. Although, it gives an probabilistic / information view of a difference in distributions. One of the outcome of KL is that it will highlight the tail of your distribution.
Euclidean distance, L2, is a metric. So it is suited when you need a metric. Also, it does not give an insight of any distribution phenomenons expect the means of the distribution.
For example, you are a teacher, you have two classes. You want to compare the grades. L2 can be a summary of how far the grades are apart. The length of tails of both grades distribution won't have impact if they have the same mean. That's good if you want to know the average level of both classes. KL will give the point of view how the class grades spread are alike. Two classes can have small KL Divergence if they have the same shapes of distributions. If your classes are very different - one is very homogeneous and the other one is very heterogeneous - then your KL will be big, even if the average are very close.