For those without a math background, the notation is very opaque. A far better explanation is to explain it numerically with simple examples.
For example, have two bits of training data:
input -> output
1 -> 0
0 -> 1
And a simple network with zero hidden nodes and train it. By hand...
Then add another bit of training data:
0.5 -> 1.5
Notice that it is now impossible to fit the training data exactly, however many training iterations we do. Now add a hidden layer with one or two nodes. Now we can perfectly fit the data, but show that depending on initialization weights we might never get there through gradient descent. Nows the time to mention different types of optimizers, momentum, etc.