Drop-in GPU Acceleration of GNU Octave
devblogs.nvidia.com
devblogs.nvidia.com
Octave is a GNU package. GNU's purpose is to ensure that you can use free software. Running to tie Octave to flashy features like GPU acceleration without first pausing to fix the initial problem of non-free GPU acceleration is putting the horse before the cart. This works against Octave's goal, to provide a free alternative to Matlab, one that lets you understand and control your computations down to the hardware level. If we don't emphasise software freedom, then there is no need for Octave, since we already have Matlab. Indeed, it is this very freedom that nvidia is abusing here to accelerate Octave's BLAS libraries, a task that would be much more difficult with Matlab, where they don't have the source code.
I know that nobody wants to even think that this problem exists and even fewer people want to fix Clover because it's such a difficult task, but it's a task that we can't ignore.
Note also that because the GPU libraries are not a system library as defined by the GPL (they are not shipped with the OS), you can't even distribute GPU-accelerated Octave object code. We consider the Octave C++ API to fall well under the domain of the GPL's copyleft.
Personally, as a GNU Octave developer, I am very unhappy that nvidia is using Octave to advertise its hardware and non-free drivers. I am also unhappy that nvidia is luring users to use non-free software, acting against our goals. I reiterate Linus Torvalds's well-known sentiments against nvidia.
The goal is "have software that works," not "have software that retains ideological purity at the cost of advancement."
the issue with free OpenCL implementations on the GPU is that the runtime will be closely coupled with the device driver, so short of open source drivers, there most likely wont be an OpenCL runtime that uses proprietary drivers.
that said, i believe both open source nvidia and radeon drivers support some level of OpenCL, although it's been a while since i've checked them..
nvidia will do what nvidia will do, they obviously have quite a lot of upstream swimming to do before they realise that CUDA is not the best approach to take.
- matlab + windows + CUDA
- octave + linux + CUDA
- octave + linux + Free OpenCL
So you're halfway there and it could be better but it could also be worse.
Writing a very high performance piece of software against a bunch of trade-secret grade hardware for which the manufacturer already provides a bunch of support for many people is not worth the trouble. As long as Nvidia supports their hardware using CUDA and OpenCL is not going to be brought up to that level this will likely continue.
A positive view would seem to me that this midway position will allow more people to move to Octave which in the longer run might free up some funds somewhere to tackle the problem you mention.
Get's mad when people use it.
(This is what Apple has done with Metal Shading Language btw, minus the open source part. And what Intel's Larrabee would have enabled.)
Even if you disregard the licensing, the union of NV + AMD + Intel capabilities (software stacks included) is so weak that it almost never worth the effort. The comically bad software from AMD and lack of hardware oomph from Intel means the only option is to require NVidia and build on their software support. This is passable for a narrow sector of activities that can tolerate vendor lock-in (Cuda), like short lived HPC projects.
All this has resulted in lack of open GPU programming languages. OpenCL is better than nothing, but even if the implementations were of usable quality, it's not a good compiler target for higher level languages and it's not a good language to write by hand.
The situation pretty much guarantees GPU computing stays in the fringes for the foreseeable future.
I believe nVidia has a similar effort going (can't remember the name), so it's still not a single agreed standard, but it's moving in that direction I feel.
However, I don't think either HSA or PTX is very important if nVidia, AMD, and Intel (wrt Xeon Phi's) don't start agreeing on standards. I haven't seen anything to convince me that nVidia wants to work within standards, and CUDA seems to have most of the market share right now, so I'm not too hopefully for portable heterogeneous computing in the near future.
To my grandparent poster: I don't think working with GPUs is that bad right now if you bind yourself to a vendor. In my experience, drinking the CUDA kool-aid (CU-aid?) isn't all that bad.
For example, the memory access used to be entirely in the fixed function hardware and the processor could only index buffers that were set up elsewhere. Nowadays it's pretty common to have full access to memory from the processor itself.
In a few years we will get GPUs that will be able to run compute tasks entirely in software (the graphics will most likely remain fixed-function for much longer) and then exposing the GPU ISA will make much more sense.
it's fragmented, true, but it's not that bad.
> If they'd only exposed the processors directly and worked to unify behind a common compiler frontend...
they don't need to, LLVM-SPIR is supposed to enable this - compile kernels to IR and let the runtime JIT it into the GPUs required binary.
> And what Intel's Larrabee would have enabled.
intel did release the larrabee, sort of. it's called the Xeon Phi (and it isn't exactly great..)
> The comically bad software from AMD
AMD have a well deserved reputation, but things aren't that bad now.
> OpenCL is better than nothing, but even if the implementations were of usable quality
i don't know what problems you run into specifically, but most of the runtimes are definitely of usable quality. ironically, it's apple have the worst runtime (but even that isn't so bad).. think of that what you will!
things have improved in the OpenCL space, and they will continue to improve for the foreseeable future. AMD and intel are doing good here, and nvidia actually do support OpenCL. questions of performance portability aside, CL is a pretty good option for people wanting to run on GPUs.
Xeon Phi is not Larrabee the GPU, it's a product of salvage & pivot from the project resulting in a HPC sidecar. (HPC sector I adressed in my original comment).
Is single-precision computations useful in the context where GNU Octave is generally used?
Most of Matlab calculations are homework or simulations. Generally speaking the real world will add more errors then double floating points will reduce.
Honestly GNU Octave has been a surprisingly fast developing product. Matlab only got GPU acceleration 2 years ago (according to my buddy who uses Matlab) (which is very fast for a GNU project). A lot of academics are picking up on it, even professors recommending it.
I think the most glowing recommendation I heard was, "Well its free, and free goes along way when you're living on a research stipend."
https://play.google.com/store/apps/details?id=jp.yhonda&hl=e...
And they are using a Tesla card in the article. It says they are using a K10.
they use a GK104 based tesla, i.e a "dual GPU" GTX680, so it has no real double precision. unless you consider 2x 95GFlops noteworthy.
1/24th (or 1/32th in maxwell) of single precision performance doesn't really sell the usefulness of CUDA.
Within the neural nets community, single precision is almost always used (at least on GPUs).
How about connecting the deep secret GPU to the no function Intel chip TSX transaction and one safe thread, while observing the side channels and effects for better GUESSING GAMES and understanding PSEUDO or fake random numbers?
Funny surprises such as?