Using LLVM to Accelerate Application Performance on GPUs
mapd.com
mapd.com
Wow, I've been reading, on Nvidia's site and elsewhere, about cuda programming on gpus and a lot of the advice involves avoiding a lot of transfers between main memory and the gpu. However, at the transfer rate quoted above, it seems like you have a device that can read all of main memory in a few seconds.
This is great ... however, does this mean all the advice about avoiding memory transfers and the programming style went with that, goes out the window?
Edit: as a beginner, I'm probably butchering the standard advice for c++/cuda programming. However, could someone summarize how things change as you get to more capable processors. A lot of the Nvidia blog posts are geared to stuff valid for all cuda version but since I'm starting fresh with a compute-intense program targeting cuda, I'd like to target the best things cuda can do.
Edit2: While the article is fascinating for the possibilities it's talking about, it's mostly a puff-piece for some (apparently) closed-source software. Anyone know an open source equivalent or some more detailed sample code for doing stuff similar to what MapD claims to do?
I wonder what they mean by "portable" in this context.
"Also, since many platforms define their ABIs in terms of C, and since LLVM is lower-level than C, front-ends currently must emit platform-specific IR in order to have the result conform to the platform ABI."
That part applies to any front end, not just C and C++ compilers.
I recall reading a while ago that NVIDIA was working on the ability to DMA directly off certain network cards into GPU RAM. Has this come to fruition?
For this kind of application, is there a meaningful speed-up by fetching data from some central data source directly into GPU memory rather than doing the network interface -> RAM -> GPGPU RAM?
In the future we may add indexes such that looking up a single or small number of rows is as fast as possible (such operations fast now since the GPU scans are so fast, but not as fast as they would be if we had indexes)
Does it accumulate values as it goes? Work on all the values together in memory?
Anyway, it seems like this sort of speed should also allow one to work with larger indices and do the sort of queries that allows.
Yes you could imagine accelerating index lookups with GPUs (I think there's some research papers on this subject already) - maybe a future project for us.