Building a Language and Compiler for Machine Learning
julialang.org
julialang.org
After a `using Flux`, julia is suddenly a machine learning language, instead of a language with a machine learning library. I'd argue it shouldn't be surprising that he found his way to Julia because Julia is one of few languages that allows one to make such packages.
His other packages such as MacroTools.jl, Lazy.jl and Zygote.jl are also well worth checking out.
On the other hand, for training and deploying models which are easily expressed in other frameworks you will find a lot more ready made pieces of infrastructure elsewhere.
julia> using Zygote
julia> fs = Dict("sin" => sin, "cos" => cos);
julia> derivative(x -> fs[readline()](x), 1.0)
cos
-0.8414709848078965
julia> -sin(1.0) # 'true' derivative in 1.0
-0.8414709848078965
So Zygote can apply AD to an anonymous function that looks up a function in a hash table from user input.I notice this in the zygote readme:
"The Julia compiler does not yet support all features needed to make Zygote fast, particularly in the presence of control flow."
Any idea of when the Julia compiler will support these features?
You don't have to hope that your AD/ML package has that function, you can just write it or find it in a package and punch it straight in. That's awesome.
Tensorflow has OP(eration)s that are differentiable that you can compose and that's it. If you want to implement something that is going to be differentiable you need to implement it using Tensorflow OPs or add your own OPs to [tensorflow] with their gradient.
With Flux you can take code written by a random guy on the internet that never thought about using his stuff in ML and Flux will be able to differentiate it anyway.
https://github.com/tensorflow/tensorflow/tree/master/tensorf...
They've also got to manually implement the Python -> autograph translation for a whole variety of language features (so any language features that get added or changed will break autograph until it's updated.
Flux gets this essentially for free, for the entire Julia language, without the need to manually build that language -> tensorflow translation layer. With the added benefit of Julia's non-trivial performance benefits.
Here's my deal. I wrote and maintain the differential equation solver library in Julia. This thing is a huge piece of code composed of almost 70 packages and took something like 5,000 commits and almost entirely pure-Julia. I will keep writing and maintaining this code as its own research project for many reasons. And, sometimes, I want to AD this code or stick it in a neural network.
It's a very non-standard application of this kind of code, so it wasn't built to do this from the start. But I am a greedy hacker. I don't want to have to re-write it onto some computational graph package (TensorFlow) or build interfaces. I want to just call some AD library function on any pure-Julia function in my package and have it output derivatives just like it was any numerical diff package. ForwardDiff.jl, ReverseDiff.jl, Flux.jl, etc. all have autodiffs that work on these routines. I find that magical. In fact, I was shocked that it can be faster than the standard way of calculating these derivatives via sensitivity analysis (we'll put a paper out on Arxiv in a few days about this). There's still a few hiccups that can be fixed up for non-ML applications, but these Julia differentiation tools have really impressed me.
This is quite shocking indeed, I'll look for the paper when it comes out. I'm particularly interested in whether this is true for the sensitivity functions used in shape optimization of composite materials. For example, when optimizing for a stiff, conductive material, eg https://www.sciencedirect.com/science/article/pii/S002076830...
That sounds really interesting, could you expand on your work/research?
Neural Ordinary Differential Equations
In particular, is this a promising approach, or do you see it as a dead end compared to generating GPU code natively? If it's promising, are there things we need to do in XLA:GPU to make it less awful for you?
(Reasons you might want to use XLA:GPU include, you don't have to reinvent all our performance and correctness hacks for cudnn, and maybe our kernels run faster since we're targeting such a limited domain?)
One thing that comes to mind here: does Julia use some kind of primitives for various things like matrix multiplication that might be difficult to export at the LLVM-IR level?
Seems to me that the primary challenge for any "next-generation" framework or language is getting people to actually use the thing. Sharing a front-end with Python and a backend with PyTorch seems like a good way to bootstrap that.
[1] https://pytorch.org/docs/master/jit.html?highlight=torchscri...
In C# you can say
Expr<Func<float,float>> sin = Math.Sin;
And then write a function Expr Derivative(Expr expr) => ...
Which will take the above sin, and compute its derivative as another Expr, which can be later compiled using Expr.Compile()In C# this has been introduced to make SQL bindings.
So far, the only difference I see is that in C# there's a distinction between expression trees and functions themselves, but in Julia there's not.
> We need a language to write differentiable algorithms, and Flux takes Julia to be this language.
Recently on HN there was some discussion of this paper by Conal Elliott on automatic differentiation in a pure functional language (Haskell): https://arxiv.org/abs/1804.00746
This is a rather large and vague question but I'm curious whether people have comments on the relative merits of Julia vs a pure functional language for supporting "differentiable programming" for ML?
And the cool thing is, Julia's wonderful generic programming facilities mentioned in the blogpost are used to make Makie generic to backend (GL, WebGL, cairo etc). Relies on this library https://github.com/JuliaPlots/AbstractPlotting.jl
Here's a cairo version of makie https://github.com/JuliaPlots/CairoMakie.jl