https://github.com/conal/talk-2018-essence-of-ad
That said, has anyone had a look at the paper? What I can't figure out is the assertion that, "In contrast to commonly used RAD implementations, the algorithms defined here involve no graphs, tapes, variables, partial derivatives, or mutation. They are inherently parallel-friendly, correct by construction, and usable directly from an existing programming language with no need for new data types or programming style, thanks to use of an AD-agnostic compiler plugin." That would be fantastic, but that's not immediately obvious to me. Generally, reverse mode is extremely non-parallel friendly because we're essentially traversing a computation graph. That traversal can be done in parallel, but figuring out how to partition that graph is probably more expensive that traversing it in serial. Now, the paper makes a big deal about not building that graph and passing along the derivative at the same time as the function. Fine. But does that method really produce a more parallel friendly code and is that as efficient as a more traditional graph or tape based method? Unless I'm missing something, I don't see any benchmarks of the method and what would be really helpful would be a ratio of the run time of a routine versus the run time of the routine with the AD method instrumented and the gradient calculated. The big question is how does this ratio hold as the method is scaled up in terms of variables. If the parallel claim holds true, then we should be able to also calculate the resulting parallel efficiencies.
Anyway, I'm genuinely curious if someone has additional information about the computational viability.