It's here: https://github.com/machine-discovery/deer :)
5 karma · joined September 25, 2023
I agree with the cubic time and quadratic space is a big limitation for now and I'm looking for ways to make them linear (or close to linear).
Adding more context from the paper. Although there is no convergence guarantee in forward calculation, the gradient computation only requires 1 iteration and always converge (see section 3.1.1), so even though the forward calculation still uses sequential method, the acceleration in backward computation might be achieved with our method.