ArrayFire, a general-purpose GPU library, goes open source
github.com
github.com
For technical questions, @pavanky is on here :-)
- Supports multiple backends, so you can run on NVIDIA GPUs, AMD GPUs, Intel Xeon Phis, and all CPUs using the same API.
- ArrayFire currently has statistics, image processing, signal processing and Linear algebra functions. We are planning to add Machine Learning and Computer Vision functions / algorithms in the near future.
- ArrayFire is a native (C/C++) library. It can be used from other languages fairly easily.
- The main goal is to make parallel programming in general (GPU programming in particular) easier and portable.
update: Downloaded repo, grepped, found externs... looking forward to playing with this :)
Here is the image.h header file for example: https://github.com/arrayfire/arrayfire/blob/devel/include/af...
extern "C" is present on line 58.
By completely open sourcing your only software product with a liberal license you seem to be turning a software (product) company into a software consultancy, is that a fair assessment?
Do you think the market conditions that lead you to this decision are very specific to GPU computing, or would you expect similar conditions in the more general scientific computing/HPC market? Would you say that it's generally simpler to earn money by doing specialized consulting than by selling technical software libraries, even though the former is less "scalable"?
If you're earning all the money with consulting and support, how do you allocate ressources to the further development of the library? Do your software engineers enjoy working on your company's product the same as working on client projects?
> If you're earning all the money with consulting and support, how do you allocate ressources to the further development of the library? Do your software engineers enjoy working on your company's product the same as working on client projects?
This is a question we have debated a lot internally. The shortest answer is that our experience building the product bring in the customers. The customer requirements can drive further development of the product.
Choosing the appropriate open source license (BSD-3 clause in this case), helps us reuse a lot of our code in a wide variety of situations.
If I understand it correctly, since you wrote the code and own the rights, you can do this regardless of the license you chose; e.g., you could have done an AGPL-3 release to the public, and continue giving license-to-use-and-modify-but-not-release to customers.
Am I misunderstanding?
BSD 3-clause on the other hand is very easy to understand and is permissive off the bat.
I bet the market conditions that led to this are broader than just GPU computing. I think it would more generally apply to any middleware business. But it is certainly more palatable for people in scientific computing and HPC to use something free and later pay for support and services and addons. Once people start really relying on the free thing, that reliance can be monetized. In this sense, it is more "scalable" to have an open source product which is readily adoptable by early users than to attempt to sell a product to buyers that have not yet started to rely upon it and have a good distance to go before reliance sets in. This is not SaaS and never will be, haha.
Allocation of future resources is something that we have considered a lot. I wrote before about opportunity costs associated with an open source business model: http://notonlyluck.com/2014/08/13/opportunity-costs-required...
We just open sourced today, but our plan is to treat the open source product the same as we have always treated it even when it was proprietary.
Great questions! Hit me up on Twitter @melonakos. Would be good to connect more with you, especially if you are going to SC'14 next week.
I do think that libraries should be distributed as open source, but I'm also hoping that at least in certain areas there is a way to commercially develop them as a product business. Provocatively speaking, if software "eats the world", then libraries are too important to just be developed as a by-product of some other ventures or in support of a platform/eco system.
Personally I'm planning on releasing a library under a GPL + commercial dual licensing scheme and later on another library under a non-commercial (incl. academic and government research) + commercial dual license. We'll see how that works out.
We need to remove mention of ArrayFire Pro from the website now. Thanks for pointing to that.
This and today's .NET announcement shows how hard it is to sell proprietary developer tools. I had considered using ArrayFire for some of my own commercial work, but in the end decided to roll my own OpenCL code in order to have better control. If you require cutting-edge performance (which is the reason you'd consider ArrayFire in the first place), there's just too much risk involved if the vendor doesn't get details like memory access order right on complex matrix problems. Open-sourcing reduces that risk quite a bit; if this decision had been made 3 years ago, I would have given the product a closer look.
From a business perspective, open-sourcing will murder their margins so they're basically gambling on their ability to jump-start volume. I think the product is in a tough position because most of the action these is going towards "Big Data," where data doesn't fit on a single machine -- let alone a GPU -- or towards heavy number-crunching, where hand-rolled kernels will outperform generic array libraries. They might have luck serving as a kind of backend to NumPy, but then they're two steps removed from the customer so it'll be hard building a relationship that leads to a sale.
As a side note, it seems odd to me that "native CPU" is a target distinct from OpenCL, which already runs on both CPUs and GPUs. I understand that kernels written for GPUs sometimes need to be rewritten for CPUs to take advantage of the different computation and memory architecture, but since their native CPU target isn't vectorized or multi-threaded, it seems like any further effort should be spent adapting the OpenCL kernels for CPU platforms rather than reinventing the wheel with a distinct C or assembler target.
I admire the general goal of making GPU processing more accessible, but it's a problem with a lot of nuance and requires a significant amount of customer education. GPUs are sort of like quantum computers in the limited sense that they're totally awesome at some tasks and totally suck at other tasks, and you need a solid grounding in the theory to distinguish the two sets of cases. Open-sourcing should at least help with the education angle, since ArrayFire now represents a respectable percentage of publicly viewable OpenCL code. (The open-source scene for OpenCL is pretty depressing right now.) In any case, good luck out there.
We are planning to move towards a single library that dynamically loads the appropriate backend depending on the runtimes / drivers available. If we completely relied on OpenCL, the same binary will not work on machines without the OpenCL SDKs installed.
> I think the product is in a tough position because most of the action these is going towards "Big Data," where data doesn't fit on a single machine -- let alone a GPU -- or towards heavy number-crunching, where hand-rolled kernels will outperform generic array libraries
Well that is two part question. As for hand-rolled kernels, they will obviously be better if you know the problem type. But more often than not, our users are happy to get "X" times the speed up in "Y" hours as opposed to "(1.2 - 1.3)X" speedup in "(3-5)Y" hours.
As for Big data, this is something we are working on / towards. We have some ideas that will make scaling across multiple GPUs and multiple machines easier. Since we will be doing this publicly, I am sure we will get a lot of valuable feedback from the community.
And you are right. Too bad we didn't do this long ago!!! Hindsight is 20-20 as they say. I wrote about some of the internal deliberations we had on this decision here: http://notonlyluck.com/2014/07/31/the-decision-to-open-sourc...
- How do you deal with software that has been previously run with coarse grained parallelism, optimized for Multicore/Multinode x86? In my experience, GPGPU porting often leads to a tedious, mostly mechanical conversion from coarse grained to fine grained, which usually includes privatizing all your data manually in your parallel domains.
- Can you do multinode / multi-GPU without wrapping everything in MPI?
- Can ArrayFire also run on CPU clusters?
- How do you deal with different storage orders? Up until now, GPUs often require a different storage order than CPUs, (wide vs. narrow vector processor) - and how does that factor into the last point?
Sidenote: I've been dealing with above problems in a Fortran based research project and have created a preprocessor framework[1] to deal with it.
The library also implements the algorithms in three backends (CUDA, OpenCL and native CPU) using the same API. We'll be adding support for SSE/AVX/NEON to make it more performance portable inthe future.
EDIT: The CUDA and OpenCL backends will obviously be faster than MKL. We'll be adding SSE / AVX support at some point which'll make the CPU backend faster as well.
We wrote some blogs on this:
-http://arrayfire.com/triangle-counting-in-graphs-on-the-gpu-...
-http://arrayfire.com/triangle-counting-in-graphs-on-the-gpu-...
We plan on looking at additional algorithms in the future.
"Real speedups for you code!"
I would not say we compete with CUDA, more like complement them!