Here's how I like to think about it:
Forwards mode AD (and finite differences) tell you how much wibble of the inputs corresponds to a given wobble in the outputs.
Reverse mode AD tells you how much wobble of the outputs corresponds to a given wibble in the inputs.
If you have more inputs than outputs (such as in optimization), it's cheaper to calculate the wibbles given a wobble, than the other way around.