Whenever I see an ML paper with surprising results I'm usually sceptical, as the results often come down to one of the following:
1. The assumptions don't generalize
2. The observations don't imply the conclusions from the paper
3. Hyper-parameter shenanigans: It's usually possible to choose hyper-parameters where method A is better than B, even if generally B is better.
In this case the lower-left part of figure 4 makes me suspicious, since forward mode seems to be better than reverse mode even on a per-iteration basis, which makes me suspect that they chose hyper-parameters where reverse mode has a disadvantage.