DARPA gives $3M to Continuum Analytics to improve Python for big data
drdobbs.com
drdobbs.com
http://continuum.io/blog/continuum-wins-darpa-xdata-funding
The money is going to develop Blaze and Bokeh:
I'm extremely excited to see Bokeh develop more.
(The Dr. Dobbs article doesn't mention licensing at all; the Continuum release implies the XDATA program has a focus on open source but doesn't specify more about Bokeh.)
BokehJS draws plots in the browser with canvas. BokehJS maintains a data model of plot models (plots, renederers of different types, tools) with Backbone JS. It has a a connection to the Bokeh server via web sockets.
The Bokeh server keeps track of plot objects and other models with redis. It keeps these persistently in documents.
The python client connects to the server and either create new plots or updates existing plots. You can create a plot with python and get a url to embed that plot in any web page. This means that your python code can just be a script that doesn't have to worry about serving up a web page.
We intend to write client libraries in other languages like java, R, and ruby to name a few.
Bokeh is starting to sound like everything I've ever wanted.
I'm still a bit skeptical Blaze will catch on. It's ambitious, and the python numeric community has had to chase each new implementation (with serious issues with backwards compatibility and cost of adoption) over the years. I would actually far prefer numpy and its limited functionality, over a much more complex system that didn't have any adopters.
You are absolutely right that both of these projects are extremely ambitious, but, given the $3 mil grant, I can't help but be optimistic.
1 http://blaze.pydata.org/docs/overview.html#blaze-is-a-genera...
You are welcome to harbor your doubts. :) There will continue to be a place for scientists who want to write software that takes advantage of the intricacies of low-level memory transfer. (Fundamentally, this is the control that C gives you.)
However, as hardware becomes more heterogenous (Xeon Phi, GPU, SSD/Hybrid drives, Fusion IO, 40gb/100gb interconnects, etc.), it's going to get harder and harder for a programmer to learn the knowledge required to truly optimize data transfer and compute within a single node or in a cluster, and still have any time left to do real science. Newer approaches to software development are needed, to maximize the potential of all this great new hardware, while also minimizing the pain and the knowledge level required of the programmer trying to utilize this hardware.
The key saving grace is that much of these hardware innovations are really geared for data-parallel problems, and it just so happens that parallel data structures are also easy for scientists and analysts to reason about. The goal of Blaze is to extend the conceptual triumphs of FORTRAN, Matlab, APL, etc. and build an efficient programming model around this data parallelism.
If you look at what's already been achieved with NumbaPro (the Python-to-CUDA compiler) in just a few months' work, I think the future is very promising.
The underlying premise for all of Continuum's work is: "Better performance through appropriate high-level abstractions".
I do not work with CFD anymore, 5 years ago I left the field and returned to Brazil, CUDA/OpenCL was just a promise back them, it was fast but no one had included the technology in their solvers.
The problem for CFD is that generally some simulations could take weeks to run, using python here would just add some more weeks just because of the overhead that the language have, if time or energy consumption is a problem people would just stay with what they have right now, for small stuff people use whatever they want (and I used Python back then, these days they use OpenFoam), for performance it will be C, Fortran or C++ compiled with the best optimizing compilers for at least a decade more. In these large problems data transportation and storage was a big problem.
Generally people use C++ these days, Fortran is only important in old codebases (although Coarray Fortran was a buzzword just like CUDA when I left the field).
Python is not inherently slower than C, C++ or Fortran. It is just a language after all. If you have a fast implementation of computational engine, like Numpy, you can reach speeds of C++ in computational tasks. If you have an optimizer that simplifies your math expressions, before running them - you can get better than naive C++ approach. An optimizer and computational engine that optimally offloads work from CPU to GPU can give an order of magnitude advantage over naive C++/blas code.
And to give you a concrete example - there is a nice Python library called Theano, that is doing just than.
Huh? People do it all the time. Naive, but not memory-leaking or improper complexity using, C/C++ etc code, beats Python hands town -- and can be more than 10-20 times faster.
>Python is not inherently slower than C, C++ or Fortran. It is just a language after all.
An interpreted language, with a not-that-good interpreter and garbage collector, non primitive integers and other such things holding it back.
>If you have a fast implementation of computational engine, like Numpy, you can reach speeds of C++ in computational tasks.
That's because the "fast implementation" is NOT written in Python.
With Python (and Theano computational engine) you can write something along the lines:
def cost(goal, prediction):
crossEntropy = -goal * log(prediction) - (1 - goal) * log(goal - prediction)
return mean(crossEntropy)
prediction = 1 / (1 + exp( .... -dot(x,w) - b + ...)) > 0.5
gradW, gradB = grad(cost(y, prediction) + (w**2).sum(), [w,b])
And then just apply that function to a matrix containing your data. That's it.
When you apply a function, it will be interpreted, converted into a computation graph, this computation graph will be optimized and parts of the computation will be offloaded to GPU with memory transfers between the host and GPU taken care of, and you will get a result in a user friendly and efficient Numpy array.The resulting computation will be nearly optimal and limited by memory bandwith, CPU - GPU bus bandwidth and GPU FOPS rate. With luck you can get close to theoretical maximum of your GPU floating point performance. And all done in a few lines of Python.
Now consider the same in C++. Yes, it can be done. But there are just no open source libraries available that can do that. Closest open-source implementation that I know of is gpumatrix, a port of C++ Eigen library to GPU. And it doesn't even come close to what is available in Python. So with C++, if you want to match the performance of these few lines of Python code, good luck studying Cuda or OpenCL and implementing the computation engine right, from the first time.
(disclaimer) I'm not in any way affiliated with OP and I actually use (and like) C/C++ a lot.
With vectorization and compiler optimizations* I doubt pure Python, well CPython running pure Python, would match Fortran and some C open source solvers that are the state of art like Lapack for dense matrices and superlu or plastix for sparse ones. It's not much because of the language, but mainly because python has more overhead than C or Fortran.
Maybe as you said python can work well in solving a single linear system using GPU. I worked with this before GPUs came to be used for that, so I can't comment on this one.
* Generally compilers optimize mathematical expressions unless they involve complicated pointer arithmetic, this is why Fortran is still used in this area, Fortran does not have this feature.
Maybe you should change buddies
The whole point is that it gives you abstractions on top of those libraries, like broadcasting, ufunc, fancy indexing, etc...
Also, Blaze is attempting to abstract the concept of having huge arrays distributed across multiple systems, clusters, clouds, whatever. So imagine having large distributed arrays that you can easily perform computations on. Could be very powerful if it works as planned.
Although scientists earn more credit here than business people, allow me to make a skeptical remark:
This is exactly the same reason they had in the past regarding with SQL. "Business people can't wait for programmers to run their queries; if they can learn an easy language, they won't have to rely on programmers to do that for them..."
In the end, the business people did not do SQL; they just force developers to do it for them, using SQL.
History repeats itself, perpetuating fallacies
Scientists also already use several easy languages (r, MATLAB, SAS, increasingly python) rather than relying on outside developers.
False. There are literally hundreds of thousands of people using SQL every day to query and analyze their data, that would not be able to write a for-loop or tell you what a pointer is. The same goes for R.
Just because the needs of business analytics have grown to a point where developers are needed to build extremely sophisticated things with SQL, does not mean that simple SQL does not serve its original purpose beautifully.
The same is true for Python. You can do incredibly sophisticated things with Python, but there are many, many scientists and data analysts that use it to do simple things that would not quite fit into a SQL mold, but also don't require a deep knowledge of CS.
('Investment' was used in the article in the general sense of "an investment in our children's future" or "investing in shared infrastructure", rather than expecting a specific liquid return like cash dividends or an appreciated exit event.)