At the end the engine that compiles the mathematical expression to hardware is what matters, and I don't think that LLVM IR that uses is the best IR for the optimizations.
At the end the engine that compiles the mathematical expression to hardware is what matters, and I don't think that LLVM IR that uses is the best IR for the optimizations.
I do not. Unfortunately, high performance and python do not go hand in hand. Yes, I know the heavy lifting is done by C/C++/Rust/Cuda/Blas/Numba and so on, but, when you run simulations for millions of steps, you end up with billions of python function calls.
Afaik only Jax actually performs any optimizations because it constructs an analytical gradient. Zygote seems to be able to that and more on the LLVM IR level which, I think should enable more optimizations.
And that something better is certainly not Julia and it's definitely not Swift.
> Yes, I know the heavy lifting is done by C/C++/Rust/Cuda/Blas/Numba and so on, but, when you run simulations for millions of steps, you end up with billions of python function calls.
For ML purposes, if your model isn't running on the GPU, it doesn't matter if you're using Swift, Rust, or whatever, your stuff is gonna be slow. Like it or not, Python is one of the best glue languages out there, and the reason why libraries like TensorFlow and Torch are not used in their native languages (C++) [0] is because they're significantly simpler to use in Python and the performance overhead is usually one function call (e.g run_inference(...)) and not billions.
If you find yourself writing a simulation, and you need to optimize away billions of function calls, you can use the C++ API provided by TensorFlow.
I prefer to just use Julia and get that for free. Also can write custom cuda kernels in pure Julia. And differentiate through arbitrary julia code, compose libraries that know nothing about each other etc
Python's performance is sufficient when the bottleneck is actually the computation done in the accelerator. In my flavour of ML, we use small models, think 3 layer NN 64 neuron wide and in some cases a small CNN. During training, most of the models reported using <150MB.
Most of the community finds python sufficient because they do not need to interleave training and simulating.
> For ML purposes, if your model isn't running on the GPU, it doesn't matter if you're using Swift, Rust, or whatever, your stuff is gonna be slow. Like it or not, Python is one of the best glue languages out there, and the reason why libraries like TensorFlow and Torch are not used in their native languages (C++) [0] is because they're significantly simpler to use in Python and the performance overhead is usually one function call (e.g run_inference(...)) and not billions.
You don't know in what I am running/doing, hence your comment comes off as ignorant.
This is the setup:
Run X number of simulated steps on CPU, collect X samples, store them in a buffer of size either X, or Z>>X, train for Y number of steps sampling from buffer, copy model to CPU, repeat for total of 1M steps, repeat 10+ times to get a good average of the performance.
All of that without any hyper parameter tuning.
Now, you also need to note that a non trivial amount of work is also done in python to augment the inputs as necessary. If the work is done in numpy, there's usually little overhead, but, it is often the case that the environment I am simulating is wrapped in wrapper functions that modify the behavior, e.g. I may need to remember the last 4 instances created by the environment, and so on. All these modifications quickly accumulate and for a single experiment of 1M steps, you end up with billions of python calls. The community uses more or less the same package/framework as the intermediary between simulations and models.
The issue is so prominent, that the community as a whole moved away from running in single core the simulations, to having multiple parallel actors to collect transitions/data points, this also requires new theory. Furthermore, there have been many proposed architectures for distributed and asynchronous training because the bottleneck is not the GPU or the model, but rather, how fast you can collect transitions. Infact, there was a distributed architecture by google that literally sends the transitions over the network into a few GPUs, the reason is that the network cost is amortized because you get to run hundreds of simulations concurrently.
IIRC, a maintainer of a popular project/framework that I contribute saw improvements upwards of 2x when using C++ over python.
That's kind of my point. Most of the community has models running on GPU's and don't care too much about CPU-bound workloads, training or otherwise.
If you do care about that, you are in a relatively small niche of machine learning, statistically speaking. I am not denying it's existence, I'm just saying that your stack will have to be different if you want to extract the maximum level of performance out of your CPU.
> IIRC, a maintainer of a popular project/framework that I contribute saw improvements upwards of 2x when using C++ over python.
That's not surprising at all. Like I mentioned in the parent, if you profiled your code and you found that Python function calls are the main bottleneck, and you believe it to be worth investing time in getting rid of them, you can use the C++ API's of Caffe/TF/PyTorch/Whatever.
I personally don't work in simulations so I haven't ran into your problem. In the deep learning world, the CPU is unusable for any task (training, inference, evaluation, etc.), so I've never been concerned with things like function call overhead.
This is becoming increasingly outdated as significant parts in scientific machine learning require parts to be written that don't simply compile to GPUs. Think, for example, when you mix a physics model, say for RF or some such, and deep learning to model parts of the function. In python, you cannot write the RF model because python is vastly too slow, so you're forced to write the RF model in something fast, like C/C++, then integrate that to python, then integrate that to your favorite tensor network, needing more languages than you can do immediately with Julia.
Deep learning is moving rapidly out of simply being a tensor engine, and being a tool in much larger problems, where many high performance pieces need developed. Julia is light years ahead of Python for these domains, and I cannot see Python ever catching up because it suffers from performance and/or multiple language problems to solve these.
If you've never learned about scientific machine learning - go read some or watch some videos. It's fascinating and growing rapidly.
I’m sure that Julia is better for some specific tasks, for example phyisical simulations, that are important for some types of scientific tasks, but I don’t see any valid argument on why a simple tensor engine in Python is not smart enough to simulate the human brain, which is just a bunch of connected neurons.
That's funny, since protein folding uses the methods I described to achieve it's results. Multiple stages in their work [1] uses physical modeling, not NNs, to do protein folding, most likely because the problem becomes currently intractable to solve with simple Python + TF. NNs are a step to adjust the physical model, exactly like I described above.
Simply read their paper and note all the physics models incorporated at about every step of the process to enforce physical constraints - this means vastly less parameters, less training time, less training data, faster evolution of the process, etc.
Here's their repo [2]. They did that work in Python and TF, likely because it was started years ago. As Julia becomes a much faster develop tool for this type of work I expect this will change. They also used TF 1.14 - showing the age of their development. They did not release the feature generation code which is a significant component; the released code only works on the specific dataset they provide. This is likely because this component is not simply simple python they wrote, but an amalgam of things written to make the physics parts of the chain fast enough. But they don't clearly state either way.
Also, by your argument, since it's possible to solve any NN problem with a network only 3 layers deep, why not just claim that's all one needs? Because it's also not computationally feasible.
The point is that by adding outside knowledge, such as physics models, you can have vastly smaller networks, require less training data, train and infer faster, with the end result of being able to solve a much larger class of problems efficiently.
So yes, you can do it in python, or any language, if you want to waste orders of magnitude more effort and resources to do it, effectively limiting the things you can practically do.
This is why Python incurs a unnecessary cost for such development.
By willfully ignoring learning about these methods and being ignorant about even the results you cited you will miss out on extremely useful knowledge.
[1] https://www.nature.com/articles/s41586-019-1923-7.epdf?autho...
[2] https://github.com/deepmind/deepmind-research/tree/master/al...
Given how divergent the current crop of ML frameworks are, is this really a realistic expectation? Having played around with Julia and Flux for ML, I find I have to do just as much rewriting when translating e.g. TF -> Flux as TF -> PyTorch. You get some limited mixing and matching with Caffe2 <-> Torch and TF <-> JAX, but that breaks down the moment you leave a company's walled garden.
> I don't think that LLVM IR that uses is the best IR for the optimizations.
I think Chris Lattner agrees, which is why he also helped start https://mlir.llvm.org/. If anything, I predict we'll see more frameworks targeting it (prototypes for Numpy, PyTorch, TF and general XLA already exist). This implies that languages that target LLVM now will actually have a leg up because their compiled semantics can be more easily lowered to something accelerator-friendly.
It may not be quite as mature, but it's getting there quickly.
It's also far more interoperable because of Julia's multiple dispatch and abstract types.
For example, the https://github.com/alan-turing-institute/MLJ.jl ML framework (sklearn on steroids), works with any table object that implements the Tables.jl interface out of the box, not just with dataframes.
That's just one example.
I don't want mature, I want to use cutting edge AI algorithms and data pipelines on the newest NVIDIA GPU, so to make me move to Julia, the researchers have to move first.
At the same time for my non-AI data processing work I'm staying with Julia.
Julia is like, the poster child for being able to mix and match and compose code and algorithms: Flux/Zygote can do auto-differentiation on the entire language without modification and you can drop in quite literally arbitrary Julia code as components of your network and it works.
> don't think that LLVM IR that uses is the best IR for the optimizations.
What makes you say this? They community has been able to do some pretty amazing things performance wise-e.g. pure-Julia implementation of BLAS/LAPACK reaching and in some cases exceeding performance parity, plus there’s been plenty of work in CUDA support for arbitrary Julia code, which is impressive.