Differentiating SSA-Form Programs in Julia (2018) [pdf]
arxiv.org
arxiv.org
The author's package, Zygote, makes all Julia code differentiable, so any program can be optimized as an ML/AI model to learn a set of parameters given some training objective.[a]
:-)
[a] That said, you will not be able magically to overcome the limits of mathematics, in case you're wondering. See darawk's and b_tterc_p's comments below.
E.g. f(x) = x + 1
Go find me what input minimizes distance to 2?
Edit: and the reference to ML is given only because that’s a common use case for non linear optimization?
Edit 2: based on parent's edit [a] I think I've got it right above. Basically this won't allow us to optimize anything that wouldn't have been possible to optimize before. But it will be a much less painful experience. Just implement it and then define the input space + target. I do feel like there will be a grey space of thinking "is this a good program for optimization?" but its a really cool idea nonetheless.
A critical feature of model fitting problems is that the data can be normalized across the training set prior to training. Without something like that, first-order methods are almost unusable as black-box methods; you have to do a ton of custom initialization and normalization work for each problem class. They don't have affine invariance although momentum methods can help. If you don't normalize your data or your model doesn't respond well to the usual forms of random weight initialization, you'll blow up or slow to a crawl on even trivial problems like linear least-squares. It's quite shocking that SGD and variants like coordinate descent work as well as they do, but it's important not to draw undue conclusions about its general applicability by understanding what properties of supervised learning problems contribute to its success.
I'm curious if ML experts reading this would strongly disagree with this general assessment. It's not my own area, but I've done a good deal of experimentation and working through things from scratch. If my conclusions are totally off, I'd want them to be corrected.
Anyway, yes, AD is valuable even if you don't use it for optimization. It has many applications in scientific computing and engineering; it was considered one of the great inventions in scientific computing long before deep learning made it mainstream. I have my own reverse-mode AD library and I recently used it for computing the constraint Jacobians for a multi-body simulator and for deriving the discrete Euler-Lagrange equations of motion for a given Lagrangian that's written as a complex expression across several coordinate systems. The Euler-Lagrange equations are then fed to a symplectic integrator to simulate the system. In a problem like that, traditionally you'd symbolically derive and then hard-code the expressions for the derivatives, but they get very gnarly to the point where you need specialized software just to derive the symbolic expressions. AD is great for rapid prototyping and experimentation in applications like that, and once things are settled you can use code generation to lock it down for efficiency and stand-alone use without the AD library.
People have done this kind of thing before, but only by e.g. reimplementing a very limited physics engine in Theano. We can just reuse existing code, along with the significant domain expertise embedded in it, and make this kind of thing a ten-line script.
[1] https://fluxml.ai/2019/02/07/what-is-differentiable-programm...
THANK YOU for doing this work and sharing it with the world. It's awesome.
Because of your work, I've started experimenting with Julia for building experimental DL/ML models, instead of Python+PyTorch, which is the predominant stack we use at work today for iterative experimentation.
(For those here who don't know, one-more-minute is Mike Innes, the author of the paper.)
I'm always happy to hear about how this stuff is being used, if people want to reach out.
People are referring to limitations regarding expressiveness.
In hindsight, perhaps my language was too fast and loose. Thanks for pointing that out.