PGStrom: GPU-accelerated PostgreSQL
wiki.postgresql.org
wiki.postgresql.org
correction: those changes to compile on macox did get pulled so https://github.com/pg-strom/devel should build with openCL on your mac to link with 9.5dev.
http://on-demand.gputechconf.com/gtc/2015/presentation/S5276...
http://www.slideshare.net/kaigai/gpgpu-accelerates-postgresq...
edit: His slideshare page http://www.slideshare.net/kaigai has even more recent presentations. I've only had time to skim much of slides, but it seems he's flip-flopped a couple of times between OpenCL and CUDA
It's actually pretty easy to have CPUs on the scale of a GPU right now. A 15 core Xeon can put out 600 gigFlops of compute so you just need 10 of them to equal one Titan X. Now, those 10 Xeons take much more silicon area and power to generate that processing power than the Xeons but that's because they have all these branch predictors and out of order execution engines and such that mean that their performance falls of much more slowly as you throw them at problems that are more complicated than DSPish code.
Although there are many companies trying to crack this many-core problem (Kalray, AppliedMicro, Adapteva, Cavium, Tilera), I'm not aware of any "successful" solution.
Of course, it was mostly slideware but we were given the promises that it would change computing as we know it, yeah right.
They really need to improve their GPU skills.
We've seen a similar thing with GPUs. The original 3DFX was this card you had to use a loopthrough video connector with because it had no 2D capabilities. Slowly it got integrated with 2D capabilities, and then grew to be GPUs, before starting to transition onto the motherboard, before eventually making its way into Intel chips. I don't believe you can buy an i or Xeon series Intel chip these days that doesn't have a GPU integrated in it. It's not the most powerful thing, but it is usable via OpenCL.
Intel currently has something called the Massive Integrated Core Architecture (https://en.wikipedia.org/wiki/Xeon_Phi) which has some really interesting technical advantages over the popular GPU approach. This started out as a separate expansion card you plugged into a PCI Express slot, but with the Knights Landing iteration due this year, it's also going to be available as an on-motherboard chip, and even support running the OS itself. (http://www.zdnet.com/article/intels-next-big-thing-knights-l... http://newsroom.intel.com/community/intel_newsroom/blog/2013...)
GPU's are an odd case because they need a lot of internal bandwidth and are minimally impacted by latency making external cards far more viable. And unlike sound cards they can easily eat more or less unlimited FLOPS.
A few of the Xeon E3 parts do have GPUs, but that's because they're targeted at workstations, apparently.
None of the Xeon E5/E7 parts have GPUs.
Take a look at [here](https://en.wikipedia.org/wiki/List_of_Intel_Xeon_microproces...) for a list.
Whule both might be turing-complete, the kind of tasks a GPU does fast are the tasks that are extremely parallelizable.
3D image creation (the most basic thing they do) is like that. So are other tasks we program them to do (e.g. running some number crunching etc).
But normal programs we run (the OS, Word, Photoshop, the web browser etc) are not like that, and we can't easily make them like that.
GPUs were designed to run thousands of threads, CPUs to run only a few of them. CPU cores do branch predictions and instructions reordering to save latency. GPU cores do not. When they need to wait, they instead pause the thread and resume some other thread on the same core.
GPUs were designed to execute the same code on many threads, CPUs to ran arbitrary code on each thread. Each CPU core has instruction fetch, instruction decode, and branching modules. For GPU cores, there’s only one fetch/decode/branch modules per many cores (32 on nVidia).
This 2 factors are among the reasons why we have 4 cores in $200 mid-range desktop CPUs, and 1024 cores in $200 mid-range desktop GPU.
- One woman can make a baby in nine months - Nine women can make nine babies in nine months - But nine women can't make a baby in one month
Each step of fetal development depends on the previous step. For a given single baby, you can't have one woman working on step #49 while another woman works on step #17.
A lot of computational tasks are similar. Think of the Fibonacci sequence: the computation of each number depends on previous results.
Reason #2: The overhead of communication and coordination. If you split up a task amongst 32 cores, those cores need to communicate with the parent process and perhaps with each other as well. This eats into your transistor budget.
Reason #3: Resource contention. It's nice that you can split up your memory bandwidth-intense task across 128 cores, but if most of those cores are just sitting around and idling because of memory bandwidth contention issues, you haven't gained anything.
Reason #4: It's hard. Programmers struggle to write parallel code. Some tools make it much easier of course.
The rest of this post is on point, though. Serialization bottlenecks (Amdahl's Law) prevent a lot of optimizations from making meaningful dents in real workloads.
The other answer to the rest of your question is probably about when the majority of software starts running entirely on the GPU.
So, let's take the Gefore GTX 980, a pretty heavyweight GPU right now. Tech specs say it has 2048 "CUDA cores" whatever that means. Well, a cuda processor is actually a SIMD core, and in this case, I think it's actually a 32 element wide SIMD core. Which makes the 980 actually contain 64 cores (Is that actually the number of cores, or just instruction schedulers)... Which is still a bit, but you can get a modern 18-core Xeon, which only puts us off in about a factor of 4 in core count.
So, the main reason for the gap at this point is functionality. The CPU does a lot of things that the GPU doesn't, and that all takes silicon that the GPU uses for extra ALU. Things like legacy instruction decode, instruction reordering, deep branch prediction pipelines, and a giant bucket of instructions that just aren't needed on the GPU, but many modern CPU targeted software really relies on for performance. And really, if the average piece of software properly took advantage of many-cores, we'd all have a ton of cores in our CPU. It much easier for CPU designers to bolt on another core than dealing with complicated branch prediction invalidation logic, or any other trick to make crappy code run well.
In all I think the gap is smaller than you think, and we will continue to see convergence, particularly as engineers start writing code better suited to the GPU, and rely less on the CPUs fancy performance tricks. But in all, there are some fundamental differences that make complete crossover tricky.
Uh, isn't this like asking "how do I do calculus without being a calculus geek"? I mean it's specialized - how do you get around being a "GPU geek" if you're coding for a GPU "at its best"?
And here are similar pojects:
I don't think that's true. A technology might only be able to become a mainstream product if the average programmer can use it, but there are a lot of niche products that most programmers will never use. That doesn't mean those technologies haven't succeeded.
If you only ever use things that are similar to the things you already know (e.g. things you can use without making errors) then you will only make very slow progress as a developer. You'll always be avoiding the things that are different to your existing skill set. Sometimes, if you want to use something radically different to what you already use, you have to put in the effort and learn something that is hard. You can't expect the technology to come to you.
technology can only succeed when the average
programmer can use it without doing much errors
GPUs have been pretty successful despite the fact that you need to be a GPU programmer to program them.Well, if you want to be successful like PHP, sure that's true.
For me, I'd sooner sue something which works well when used by a person who's taken the time to learn what they're doing. It's not important to me what a person of median skill and effort can accomplish: there's only so many things the average engineer is going to be able to learn, for whatever reason.
You must complain a lot.
https://wiki.postgresql.org/images/5/50/PGStrom_Fig_MicroBen...
Or am I reading that incorrectly?
How is that possible?
I'll guess:
It's all in-memory, so not I/O-bound. Rows are processed in parallel with PGStrom (in 15MB chunks), while Postgres does it sequentially in one thread. Apparently moving the data around is fast enough not to matter.
Many interesting queries in normalised databases involve a lot of joins. In a BCNF database most interesting queries use 3 or more joins. Whether or not I believe the microbenchmarks is another thing :)
The gpu was designed so you can, lets say, move all vertices in a character's body 1 inch forward, you would need to process this one translation for potentially thousands of vertices. It was also designed to do bitmap operations very quickly where you need to get the same image, but smaller so the GPU would combine nearby pixels to provide a good heuristic of what that same image would look like further away.
That said if you don't mind just reading diffs you can get an idea of the state of the project by taking a gander at their github: https://github.com/pg-strom/devel
(Source: I've done a fair bit of experimentation in GPU queries in the domain of CSS selector matching.)