Yes, neural networks are objects that compute; there's even this "universal approximator" theorem that says a basic, albeit sufficiently large, neural network can approximate any arbitrary function (from a broad class of functions) to arbitrary precision. However, the theorem says nothing about whether you'll ever actually _find_ the neural network that corresponds to that function. This is what training is for, it allows us to find (the parameters of) the NN that we want to do some computation.
In other words, training is how we program NNs, but in general it can be really hard to arrive at the "program" you're looking for.
Indeed! Most of the time computers are computing, not being programmed. Yet most of the time, neural networks are being trained instead of computing.
That was exactly my point.
But also note that what you say is not necessarily true that the NN's spend most of their time training. Maybe you've got to spend a week on a huge GPU cluster training some autonomous-driving algorithm, but then it runs in "compute" mode for hours a day in tens of thousands of cars.
This practice is well known but here's a concrete source from Andrej Karpathy:
https://karpathy.github.io/2019/04/25/recipe/
"init well. Initialize the final layer weights correctly. E.g. if you are regressing some values that have a mean of 50 then initialize the final bias to 50. If you have an imbalanced dataset of a ratio 1:10 of positives:negatives, set the bias on your logits such that your network predicts probability of 0.1 at initialization. Setting these correctly will speed up convergence and eliminate “hockey stick” loss curves where in the first few iteration your network is basically just learning the bias."