Introduction to CUDA C
infoq.com
infoq.com
However, the main difference is the number of different types memories there are available. That makes GPU programming very tricky indeed. There's around 6 different types of memory visible to the programmer, each with it's own distinct access times and restrictions. On the host side (CPU), there's regular memory and page-aligned DMA buffers. On the device (GPU), there's constant, global, local and texture memory. Most of the time it's clear which memory you should use, but occasionally it takes some thinking. Especially deciding whether something should go to texture memory or global memory can be difficult without doing it first and benchmarking.
As for memories, it becomes natural once you start programming CUDA day in and day out and understand the algorithm well enough :)
generally (and especially in the case of the upcoming GK110 chip), you should use global memory. GK110 improves this with LDG, which allows you to get some caching benefits of texture (spatial locality) without having to jump through the API hoops required to use textures.
(full disclosure: I run the CUDA driver team at NVIDIA)
Oh, cool! This opens doors for applying GPGPU to a whole new class of algorithms. I clearly must update my GPU knowledge.
Global mem vs. texture mem is always the biggest choice. Sometimes it's worth thinking whether there's a potential win in caching texture/global memory fetch results in local memory. So there's still choices to be made, even though the hardware has become better and easier to program.
I used different (OpenCL-like) terminology, but we're talking about the same things. Shared memory <=> local memory. Pinned memory <=> page locked dma buffers.
and i am not sure "page aligned" and "page-pinned" are the same thing.
Makes me think that there are enough people working on GPUs out there, just not enough to stand out from the web-dev or other related news that usually get voted up..
And I'd be interested in seeing who's using GPUs for science!
http://www.accelereyes.com/arrayfire/python/
Its a freemium model, should cost you next to nothing to try (as long as you have a gpu).
Also I see that the free edition tries to connect to an outside server to a high port, and that would work here. Our network is not really friendly: I cannot even check out a git repo with the git protocol.
I can see the appeal of your product though. Some of my colleagues are trying to use PyCUDA but have to learn C in the process.
Here is a micropolygon rasterizer I wrote in OpenCL:
https://github.com/ginkgo/micropolis http://www.youtube.com/watch?v=09ozb1ttgmA
Its performance is competitive with the OpenGL hardware rasterizer, because rasterizing small polygons tends to get very inefficient in current hardware.
GPU computing is promising, but given that it is fairly difficult to predict how much of a speed-up to expect before the actual work is done, I find myself asking whether some of the expended effort is really worth the trouble. It seems that there are two classes of applications where this makes sense:
1) Minimising latency/increasing responsiveness for smaller algorithms such as in a user application or service.
2) Doing large volumes of computation in a high-throughput system.
At the moment, the CUDA platform is well ahead of OpenCL in terms of maturity, features, tools and documentation. However, OpenCL runs on CPUs, both NVIDIA and AMD GPUs and work is being done towards targeting FPGAs [1]. Interestingly, Clang also supports compiling OpenCL kernels directly to native code.
In all likelihood these platforms will stay outside the mainstream until better abstraction layers exist to shield programmers from the low-level architectural details without sacrificing performance. Something similar to the directive-based OpenACC standard that is able to perform close to hand-optimised close would go a long way.
Having embarrassingly parallel problems also helps...
-- This means, that the devleopment time will be more because you are going to be writing your own primitive functions which are available as a myriad of libraries in CUDA.
2) OpenCL can run on cpus, gpus and more exotic hardware like FPGAs, Cell Processors.
-- On Large clusters, OpenCL's heterogenous capabilities make it more appealing as opposed to CUDA which will essentially require you to write and maintain two code bases (one in CUDA, the other using pthreads or what have you).
3) OpenCL spec is designed by committee (the Khronos group) as opposed to CUDA (developed by NVIDIA).
-- This means CUDA iterates much more quickly than OpenCL. They also have control over the hardware meaning they can introduce some hardware optimizations that may not become standardized in OpenCL.
4) NVIDIA has opened up CUDA a little, implementations of OpenCL remain closed.
-- NVIDIA has thrust which is kind of open source, but their CUBLAS and CUFFT libraries are closed. They recently Open sourced part of their NVCC compiler and hooked it up with LLVM. The OpenCL impelmentations are vendor specific and AFAIK none of them are open.
--------------
Sorry if I am not coherent. Its pretty late and I had to type it forcing myself to stay awake :)
I would be really interested in these results if you can share them with me (contact@pavanky.com).
Are there options for getting CUDA going on an AMD/ATI chipset?
When it comes to the basics, CUDA and OpenCL are pretty much identical to each other for most parts. There are subtle differences in the API.
When you go deeper, there are a few more differences which mostly come from the hardware. Programming two different devices with OpenCL can be very different if you want to write efficient code. The difference between programming for different pieces of hardware is much larger than the difference between the API's used.
As a single vendor effort, CUDA moves forwards faster and has support for more of the latest GPU features. But it only runs on hardware from that single vendor.