Numba: High-Performance Python with CUDA Acceleration
devblogs.nvidia.com
devblogs.nvidia.com
It is a python wrapper around https://github.com/arrayfire/arrayfire and allows the code you write to use CUDA, OpenCL or x86.
It is still a work in progress and requires some upstream changes in arrayfire to support the numpy api better.
EDIT: Let me hedge that a bit, to advanced AVX instructions, as LLVM can do simple loops and such.
http://releases.llvm.org/4.0.0/docs/ReleaseNotes.html
Really impressive how many new things came to LLVM this year!
Edit: Oh, sorry you meant that other guy's link to LLVM's vectorization tutorial. Ignore my reply ...
I was going to say maybe it's a new thing, but the following post also talks about SSE4 and is from 2011: http://blog.llvm.org/2011/12/llvm-31-vector-changes.html
Maybe it only supports a subset of SSE4? Do you know the details, compared to other compilers?
EDIT: Sorry, to reply to your question, my concern is not GCC vs. clang; If you want max out your vector ops, I would suggest you should compare to ICC as the "standard", at least on Intel CPUs.
This was on an AVX2 machine, testing the (auto)-vectorization performance of expression templates. Anything with a compile-time unknown stride or a random gather failed horribly with both. Using (semi-)explicit vectorization turned out to be much faster still.
Disclaimer: by "vectorization" I'm referring to SSE4, AVX and AVX2, I haven't had a chance to try out AVX512 yet.
For almost all practical application, I use pandas or keras / tensorflow. I'm probably biased as I mostly work with simple data that doesn't require complicated calculations.
Would somebody have some benchmarks against pandas for some standard operations ?
question: how do you debug a function which has a super complex decorator slapped on top of it?
With difficulty. If you find yourself in a situation where you have a function that works without the numba decorator, but fails with it then it's time to break out your llvm reading skills: http://numba.pydata.org/numba-doc/0.10/annotate.html
That being said it is very rare that that happens.
I believe Scipyconf 2016 had a talk on numba where he goes into it in great detail. Just search it up on Youtube.
Anything that is not convenient to be written as numpy arrays can be written using numba. Also it works with pure python code so your prototype can be used at scale with nothing but a decorator.
import numpy as np
arr = np.array([1, 3, 7, 5, 4, 3, 1, .100])
maxval = arr.max() # 1st pass
minval = arr.min() # 2nd pass
Whereas with numba you'd have something like this: from numba import jit
@jit
def maxmin(arr):
maxval, minval = arr[0]
for e in arr:
if e < minval: minval = e
if e > maxval: maxval = e
return minval, maxval
And that will get optimized to numpy-like speeds, but with a single pass over data. So for large arrays, you'll get about 2x speedup, since memory access is the bottleneck.As for optimizing this use case for NumPy, I'd go for a cythonized maxmin() function. Which is pretty much the same numba does, but you're moving the compilation overhead from the JIT into the compiling step of the module.
Regarding the moving the calculation, yeah. I get that. I argue that the compiling step of the module happens once for the module, no? The JIT will be something you force onto every execution. Right?
And none of this actually means this library shouldn't have been made. Just that it is a poor example for why it is better.
Though, I would also think a general .stats method would make the most sense. It is quite often nowdays to want a lot of stats.
And again, I am not arguing that the new lib shouldn't exist. I just question that example.
I'd add some "intelligence" (if you will) to the class, so that when I calculate the max(), I'll also track the min, average, etc, save those values in an internal cache, and return the max. Next function call (wether it's max, min, avg, or any other of the cached values), it gets the result from the precomputed cache.
Of course, some operations will affect said results, and there are two ways to handle that: either modify the cached results, if we're talking about some change that can be formulated (say, multiplying by 2 will multiply every computed value by 2), or simply invalidate the cache and re-compute the next time one of this functions is called.
As it was said, this would mean calculating some things that aren't the explicitly asked thing (max, min, avg, etc), and giving a worse case scenario, in the hope that it will result in a speedup on some usecases.
My guess as to why they don't do it? Because it's not a "generalizable" problem, and you have the machinery at hand to implement your custom function that fits your particular use case (say, using Cython and the numpy/cython wrappers).
It involved tons of time series data that followed a state machine, with very little training data.
A useful algorithm to force a series of noisy predictions to follow a state machine is the Viterbi decoder.
Numba let me write a JITted version that got order of magnitude improvements, especially when there were over 10^8 time series points.
It's a great piece of software, if a bit finicky sometimes.
pandas creator here. Numba is a complementary technology to pandas, so you can and should use them together. It is designed for use with NumPy arrays and so does not deal with missing data and other things that pandas does. It does not help as much with non-numeric data types.
Depending on what you're doing, you might be able to use numpy.nan as a value. It does work inside of numpy arrays. But some methods that operate on those objects might not work as you expect.
For instance, if you run numpy.mean on a numpy array of [nan, 4, 5], it will return nan. If you run the same thing on a pandas dataframe of the same values, you'll get 4.5.
I was confused by your comment though, specifically the idea that you are using tensorflow because you 'mostly work with simple data that doesn't require complicated calculations'. This seems very contradictory. Have I misunderstood?
Numba can give you similar performance to what you'll see for highly optimized operations in pandas (e.g., for groupby-sum or a moving average), but you would have to write the low-level loops yourself. Like Wes writes, it's a complementary technology: in principle, the low-level loops in pandas currently written in Cython could be ported to Numba instead.
I had a little project where I experimented with this a few years ago: https://github.com/shoyer/numbagg. Since then, I expect numba performance has only improved.
Note that this is unlikely to ever happen in pandas itself, for various reasons. The existing routines in Cython already work (and unlike Numba, have a good story for distribution), and the algorithmic core for next version of pandas is being written in C++.
The thing where the Julia CUDA support really shines is that it is supporting arbitrary Julia structs and not just a blessed few datatypes like Float32.
Note, this post was originally published September 19, 2013. It was updated on September 19, 2017.
Continuum gradually open sourced all of it (and changed their name to Anaconda). The compiler functionality is all open source within Numba. Most recently they released the CUDA library wrappers in a new open source package called pyculib.
Some other minor things changed, such as what you need to import. Also, the autojit and cudajit functionality is a bit better at type inference, so you don't have to annotate all the types to get it to compile.
We thought it was a good idea to update the post in light of all the changes.