You could view it as 'thousand brains' with lots of weight sharing for efficiency+regularization.
- there's not a "correct" training prediction
- there's not a "final layer"
It's about computing the amount by which you adjust the weight.
And unless there's been a major development in neuroscience that I'm not aware of, backprop is not the way the brain does it.
There is no global “teaching signal” or “delta rule” error correction. Learning via reward and punishment is a wrong level of abstraction for fundamental cognitive tasks like visual object recognition or “parsing” auditory signal.
Sometime people mention dopamine as a kind of reinforcement signal but it operates on completely different time scale, orders of magnitude slower than any iterative optimization model would require.
And the energy and time spent on iterative optimization in ANNs is not available to living organisms with constrained resources.
If you’re interested in authoritative opinion on what kind of learning is biologically plausible see e.g. prof. Edmund Rolls recent book called “Brain Computations”.