Postgres on the GPU
wiki.postgresql.org
wiki.postgresql.org
On the surface, this seems very similar to https://news.ycombinator.com/item?id=5592886 except it's nicely integrated with Postgres. Or am I missing something?
tmostak?
That seems to be pretty much it, although the read-only part is a limitation of Postgres rather than a feature of the module.
Furthermore, it seems to be dispatching queries intelligently so you can perform all queries against the FDW table with (I hope) minimal overhead, if the qualifier can't be compiled to a kernel the fdw will run it as a normal on-CPU qualifier. That's a thoughtful touch.
Alas, Tim's been talking about this for two years now, and as far as we know he's the only one whose ever seen the code.
Anyone can explain why opensource projects embrace CUDA over OpenCL? As I understand OpenCL is more generic API which could be potentially used with CPUs and GPUs.
Firstly, CUDA is just more mature; there is a very large and well-established set of libraries for a lot of common operations, there is a decent sized community, and Nvidia even produces specialized hardware (Tesla cards) designed just for CUDA.
Second, all that generic-ness of OpenCL doesn't come for free. With Nvidia, you're just working with one architecture; CUDA cards. Optimizing your kernels is much easier. OpenCL is just generically parallel, so you could have any sort of crazy heterogeneous high-performance computing environment you have to fiddle with (any number of CPU's with different chipsets and any number of GPU's with different chipsets).
I haven't used OpenCL myself, but almost purely anecdotally I have heard many people say that CUDA is often slightly faster[1] and the code is easier to write.
TL;DR: CUDA sacrifices flexibility for ease of development and performance gains. OpenCL wants to be everything for everyone, and comes with the typical burdens.
[1]: Maybe this is a result of OpenCL being more generic and so harder to optimize.
I utilised the OpenCL programming interface to write code that would run the same kernel functions on CPU and/or GPU devices (using heuristics to trade-off latency/throughput) which is something that is not possible afaik using the CUDA toolchain.
TL;DR YMMV and horses for courses.
Since I am working on code generation of Kernels to perform dynamic tasks, I can't afford to write at the lowest level available. (I'm accelerating Python/Ruby routines though so OpenCL gives a significant bonus without much pain at all.)
[1] http://dl.acm.org/citation.cfm?id=2066955 (Sorry about the paywall, I access through University VPN)
Nvidia is in the slow process of eventually discontinuing further CUDA support, and it is recommended to write new code in OpenCL only.
[Citation needed]
Their OpenCL support is still limited to v1.1 (released in 2010), while just few months ago they've released a new major version of CUDA with tons of features nowhere to be seen in (any vendor's) OpenCL.
[1] http://www.techpowerup.com/181585/NVIDIA-CUDA-Gets-Python-Su...
[2] http://www.mathworks.com/discovery/matlab-gpu.html
[3] https://www.quantalea.net/media/pdf/2012-11-29_Zurich_FSharp...
Well that's certainly not true in the general case.
"Because it's hard" is a cop-out.
"It's too hard to accomplish given constraint [X]" where X is a deadline, financial constraints, or other real/tangible resource limitations might be one thing. But if you're working on your own timeline on some sort of open-source project, or there is nothing external preventing you from acquiring the expertise/resources to conquer the hard problem, then "Because it's hard" is an absolutely shitty excuse to not do something.
That said, if it makes sense for your project, make it happen! :)
I'm not trying to convince you you don't need it or shouldn't do it, I was looking for a datapoint about what you find valuable in OpenCL.
NVidia seems to be the preferred hardware for institutions/big companies. I'm not sure if this is because NVidia's architecture is better for supercomputers or if they're simply better at marketing to those types of customers
NVidia funds a lot of academics in my space, and I've found academia to be very anti-open source for those reasons, which amuses me greatly.
Case in point, Matlab. Why is this taught in a world with Python/Numpy/MatplotLib?
In my area of CS (artificial intelligence) it seems considerably less popular. I don't really remember how to use it, since the last time I used it seriously was in some engineering (but not CS) courses in undergrad.
You may be right regarding Professors, but that is also changing as they age out.
In the world of electrical engineering Matlab can do things that other packages can't.
From personal experience, I have spent many hours looking at the results of a atan2 function in C++ and Matlab and trying to get them to agree. After a day of work I was able to get them to agree by precisely controlling the rounding modes and using my own atan2 function. This was not fun, and I would rather give somebody 1k to take care of it for me.
Wait, are you saying academia is anti-opensource because of nVidia funding?
I don't see a serious problem with the scenario you describe, though. You're not really describing a hostile scenario, just an affinity for commercial software.
There is a bit of a problem of course; I find a lot of papers that describe how to do things with commercial technology that isn't in the budget. That hasn't been insurmountable for me in any way, but maybe others have had more serious problems with it.
[1]: http://sagemath.org/
The advantage of OpenCL is that it runs on more platforms (not just NVIDIA). The problem is that it's more complicated and more of a headache.
My advice to programmers is to start with CUDA and play around with your problem for a while. Time spent learning how GPUs work, what kinds of operations are efficient, and how to parallelize algorithms is not wasted if you switch to OpenCL later. Once you've made some progress then make an informed decision about whether you want to go to production with CUDA or OpenCL.
Because in bitcoin mining, it's always ATI/AMD and OpenCL because ATI cards are like ten times faster than Nvidia cards. This is because of architecture differencies.
Does it not affect this postgres table-scanning task? I wonder if they did any benchmarks.
http://www.extremetech.com/computing/153467-amd-destroys-nvi...
False. I authored a Bitcoin miner utilizing this quirk (bit_align). I was also the first to leverage another instruction exclusive to AMD (bfi_int): https://bitcointalk.org/?topic=2949 bit_align "only" gave AMD a 1.7x advantage over Nvidia. The biggest perf gains (2x-3x!) came from the fact AMD has more execution units: https://en.bitcoin.it/wiki/Why_a_GPU_mines_faster_than_a_CPU... (I also authored this section of the wiki).
https://en.bitcoin.it/wiki/Mining_hardware_comparison
The fastest Nvidia showing is the Tesla S2070 which is a $18k Server with 8 GPU's! It can just barely keep up with a slightly over clocked single gpu HD7970.
So is nvidia. Both companies produce absolutely horrible drivers. Not just for linux either, tons of bluescreens, crashes and other windows instability issues are video driver bugs. That is what happens when the sole concern is speed, and stability is totally ignored.
Perhaps it's just my scars showing, but I'm concerned about database system stability with active GPU hardware added to the box.
I probably wouldn't add the GPU hardware to a master but rather do the queries that would benefit from it on a streaming replica, where an occasional kernel panic won't be so severe.
Regardless, looking forward to trying it.
Ignoring Windows as I guess I don't really take Windows servers too seriously.
I've toyed with the idea of doing pattern matching (and graph rewriting) on the GPU before but this looks like it's much more advanced than I thought was feasible.
I'm surprised they went with CUDA instead of OpenCL though. CUDA is proprietary NVidea technology and does not work for non-NVidea devices.
[1] http://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorith...
The long term trend is for the GPU to merge with the CPU, so I think we'll see more of this in the future.
Its written in CUDA, not OpenCL. Stop that.