The cost of Dual numbers (a form of forward-mode differentiation) scales linearly with the number of derivatives (just like finite differencing, but more accurate). Backpropagation, or reverse-mode differentiation, is a constant factor times the cost of a function evaluation. For neural nets with millions of parameters, backpropagation is going to be millions of times faster than dual numbers.
This guy did neural nets with dual numbers in julia and found it to take ~1.2 times as long as standard backpropagation, which suggests that if you had dedicated hardware for it it would be very nice.
Dual numbers are just the wrong approach for neural nets, which have many parameters. The amazing thing about backprop / reverse-mode differentiation is that you get the derivatives wrt all the weights in one reverse sweep through the network.