Fastest autograd in the West
arogozhnikov.github.io
arogozhnikov.github.io
https://github.com/casadi/casadi https://github.com/mitsuba-renderer/drjit
DrJit is made by same author of pybind11 and nanobind.
you can't really implement it in a library in a normal programming language, because it has to introspect on how your program is computing everything, which is not something that a library can normally do; you have to implement it in a programming language, at least an embedded domain-specific language but in some cases something like a fortran compiler
That's how pretty much how all python DL frameworks are built. And thats always why you can't call .numpy() on a torch tensor which requires a gradient. Doing so would break the computational graph.
this is less true for forward-mode automatic differentiation, which you can do a pretty reasonable job of in python as just a dual-number class which you can pass to many, though not all, ordinary numerical computation functions. unfortunately forward-mode automatic differentiation is unusably slow for gradients with respect to large parameter vectors
> you can't really implement it in a library in a normal programming [...] you have to implement it in a programming language, at least an embedded domain-specific language
So which is it?
> A DomainSpecificLanguage that is defined as a library for a generic "host" programming language. The embedded DSL inherits the generic language constructs of its host language https://wiki.c2.com/?EmbeddedDomainSpecificLanguage
https://github.com/sradc/SmallPebble
The first line of the description is:
> SmallPebble is a minimal automatic differentiation and deep learning library written from scratch in Python, using NumPy/CuPy.
The author of this library wrote a blog post, which is in my opinion the best introduction to reverse mode automatic differentiation that exists:- it's very much taking the edsl approach
- it's only implementing forward-mode
As for the edsl, you can claim it is, but why? There is no parser, no syntax, nothing that would make any user think they are learning a new language. Would you call numpy an edsl ?
embedded dsls don't have parsers; that's what makes them embedded. numpy is solidly in the center of the embedded dsl concept
I see. It's actually worse than both forward and reverse mode.
Let's say you have the simple program x1+x2+x3+x4+x5. Internally the program calculates x6=x1+x2, x7 = x6+x3, x8=x7+x4 and result = x9 = x8+x5. In forward mode you get all the partial derivatives for all intermediate variables w.r.t the inputs x1 through x5, so you get things like dx8/dx2, but you never calculate (or store in memory) dx8/dx6. In reverse mode you get the derivatives of the result w.r.t. all intermediate variables, so things like dx9/dx8, dx9/dx7, etc, but again you never need to calculate dx8/dx6. In this library you calculate recursively all dxj/dxi and store them, as long as xj has a direct or indirect dependency on xi. This is a strict superset of the derivatives calculated in both forward and reverse mode.
Good catch.
AFAIK autograd generically refers to a computational graph's gradient being automatically derived from those of it's constituent computational nodes/operators. It's not limited to cases like PyTorch that presents itself as "differentiable programming" where the graph is implicitly built by code execution.
The alternative to a differentiable language is something like the original Lua-based Torch, or TensorFlow's original non-eager mode, where the computational graph is explicitly built by library calls out of predefined nodes/operators, each of which knows it's own gradient.
https://jax.readthedocs.io/en/latest/faq.html#jit-decorated-...
Now I'm curious to know where that is useful. Some kind of meta-learning approach?
I imagine any kind of trajectory planning for many agents faces similar challenges (e.g. robots in a factory).
For now I'm testing one-after-another because failure of previous (overflow, too-small-transfer, etc) gives a hint about which part can be changed in the graph.
I've been thinking about vectorizing over multiple guesses at once, but don't see a fast reliable way to merge several graphs (also that induces significant burden on other parts of the project - autograd is just one component).
I recommend trying out TCC, which compiles C very fast.
Another commenter mentioned Dr.Jit which seems to be designed for this use case. This is a quote from their project page.
> Why did we create Dr.Jit, when dynamic derivative compilation is already possible using Python-based ML frameworks like JAX, Tensorflow, and PyTorch along with backends like XLA and TorchScript?
> The reason is related to the typical workloads: machine learning involves small-ish computation graphs that are, however, made of arithmetically intense operations like convolutions, matrix multiplications, etc. The application motivating Dr.Jit (differentiable rendering) creates giant and messy computation graphs consisting of 100K to millions of "trivial" nodes (elementary arithmetic operations). In our experience, ML compilation backends use internal representations and optimization passes that are too rich for this type of input, causing them to crash or time out during compilation. If you have encountered such issues, you may find Dr.Jit useful.