AMD releases APPML source code, creates clMath library
developer.amd.com
developer.amd.com
1. http://developer.amd.com/tools-and-sdks/heterogeneous-comput...
Of late, I was thinking that AMD has not been doing well in the CPU division and now I feel that there is hope and we may see them pull off another Athlon...
It is in very active development and the community is very nice and helpful.
I think there also doesn't exists something similar, i.e. a lib which can easily do the calculations on both the CPU and the GPU (via OpenCL).
In my own library Aura, I focus on maximizing performance (over developer convenience). The target audience are developers of real-time applications that need every last drop of performance from their hardware, while still maintaining a sane and cross-platform API. I strive for a Boost.Asio for accelerator developers. Aura has a already a rudimentary wrapper for clFFT, clBLAS is in the works. So the idea is to, for each platform, utilize optimal vendor-supplied library functions and combine them in a coherent interface.
[0] https://github.com/kylelutz/compute
[1] https://github.com/ddemidov/vexcl
Atm., I want to re-implement some of the functionality of Theano but in C++. Esp. I also want to implement deep neural networks. And I want my code in a way so that I can easily switch between different calculation backends later on, like using the CPU or multiple CPUs (hopefully with vector processing), or some GPUs (e.g. via OpenCL) or even some multiple-machine cluster.
I found many libs doing one of this very well but only very few libs which supports multiple backends like ViennaCL. For example, Boost.Compute only supports OpenCL but not the CPU. VexCL, as far as I understand, also does not support CPU calculations.
In what state is Aura? And would it be a good fit for my needs?
CPU is supported through OpenCL by these libraries. Remember, OpenCL code can run on CPUs. As for Aura, it is pre-alpha, not a lot of functionality there yet, I'm still figuring out the interface. So not usable yet, but keep an eye on it, it will be.
For me, the most important thing in both CUDA and OpenCL is the programming model. It allows us to describe data parallel problems and related data (in)dependence explicitly. Compilers should be (and already are) able to generate efficient code from this. It is not as nice as it could be, we have to write kernels by hand etc. But there are libraries that make our lives easier (we discussed them earlier). And there is also C++ AMP which tries to integrate better. Yet still, while we have all these options and we can solve most of our problems with more or less effort and elegance, I believe there must be something better out there: the right way to describe data parallel and task parallel problems as well as concurrency etc. Maybe the FP guys are on to something, I don't know. I'll be on the lookout.
[0] http://www.pds.ewi.tudelft.nl/fileadmin/pds/homepages/shenji... [1] http://comparch.gatech.edu/hparch/papers/lee_plc2013.pdf
it's still (practically) impossible to do a proper LINPACK benchmark with open source tools on AMD GPUs, although this is a step in the right direction, and more importantly, a big blow to CUDA.
Unfortunately, I'm pretty sure NV is entrenched at this point. My AMD card and my resolution to port everything I needed to use lasted about 3 months, and that was without any CUBLAS dependencies :(