GPGPU Accelerates PostgreSQL
slideshare.net
slideshare.net
If any of the project team are reading, what I'd like to see most is GPU-accelerated point-in-polygon lookups in postGIS, ST_Contains and so forth.
And let's admit it, writing SQL statements with `vector group by ...` make you feel like a bad-ass.
Often the best way is to divide polygons into tiles that have a limited number of verticies (this also makes indexing much more effective).
Edit: right, after thinking about it, the branches can be optimized out. It could be fast if there are a set of sorted segments, just parallel compares and some boolean logic.
Which leaves the problem of getting the data to GPU. Because you can definitely stream same comparisons on the CPU much faster (memory bandwidth limited) than you can stream the data to GPU over PCIe.
So 2 CPU socket system, such as Xeon E5, I'd bet on the CPU. PCIe 4.0, 16 lanes would give 30 GB/s (not sure if PCIe 4.0 is supported anywhere), vs. aggregate CPU bandwidth of up to 150-200 GB/s. Dual socket Xeon E5 supports at least 1 TB of RAM (16x 16 GB buffered DDR4). 32 GB DDR4 memory modules exist as well, I think, and larger DIMM banks than 16 slots can be supported. It's just the number of slots on typical 2 socket mainboards.
With more realistic setting, CPU would be even more ahead. All of this is ignoring GPU latency issues, which can be anywhere from microseconds to tens of milliseconds in pathological cases.
Unless the data was on the GPU in the first place... I think currently a single GPU can have up to 12 GB of RAM. Maybe larger GPUs exist too. That's just not much RAM compared to what is typical for CPUs. Currently smallest amount of RAM a dual socket Xeon E5v3 standard server can have is 64 GB, if all memory channels have at least one DIMM.
However, point-in-polygon is a fairly simple algorithm, and if each polygon was mostly < 40 vertices, I suspect a GPU might be faster. However, for more complex algorithms, GPUs don't do as well and with many more vertices I suspect GPUs won't do as well for point-in-polygon tests.
In terms of raw theoretical FP processing power, GPUs look good - but when you start to do more complex things with them - i.e. when branching happens a lot, say with path tracing, they don't look as good. E.g. a dual Xeon 3.5 Ghz quad i7 (costing ~£950 each) is as fast at path tracing as a single NVidia K6000 costing ~£4100.
Pragmatically, it produces results of similar quality quicker than CPU-based competitors.
It is really taking the high end rendering world by storm this year.
That's a biased renderer that uses all sorts of caching and approximations that there's no CPU-based renderer that supports (VRay's closest with it's ability to configure primary and secondary rays using different irradiance cache methods), and as Redshift doesn't support CPU rendering, it's hardly a comparison worth talking about as you'd be comparing different algorithms. The pure brute-force without any caching numbers I've seen for it don't look any better than the other top CPU renderers doing brute-force MC integration.
Also, a quibble, but I guess by "high-end rendering world" you mean archviz (where VRay and 3DSMax are dominant) and a few small VFX studios who happen to be running Windows?
Really?
I know companies like Blur are trialling it, but they're still using VRay. I know The Mill have done stuff with it, but they're still using Arnold too.
There is limited software support right now, because this architecture is very new, but on the benchmarks that take advantage of the on-die GPU, AMD's latest can keep up with and surpass much more expensive i7s. We're still at a point where it's unclear that AMD's HSA will take a commanding lead, but it's promising, especially considering the price/power requirements for an A10 (~$160 currently), vs the equivalent of a high end GPU and a Xeon.
You could, ignoring storage and peripherals (reasonable in a server farm arrangement) put together many more iGPU boxes than Xeon/dGPU boxes.
Even assuming moderate gains from GPU acceleration, high-throughput database servers could be made cheaper through this method.
for streaming data into a CPU you'll be lucky to get double digit bandwidths. peak figures are ~50gb/socket, but for anything more than a memcpy, it drops off like a cliff. then you also have NUMA issues, bank conflicts, TLB misses if your data is big enough..
i've written codes that sustain >270gb/s on high end GPUs - it's not trivial, but it can be done.
you are correct though, about the quantity of GPU memory available on an average GPU. the AMD S9150 has 16gb of ram. very high for a GPU, but nothing compared to high end servers.
> PCIe 4.0, 16 lanes would give 30 GB/s
afaik it's not in anything, so we're limited to 6gb/s for GPU <-> host.. :/
> With more realistic setting, CPU would be even more ahead.
depends. getting a good fraction of peak bandwidth on a GPU is fairly straightforward - coalesce accesses. some algorithms need to be.. "massaged" into performing reads/writes like this, but in my experience, a large portion of them can be.
getting a decent fraction of peak on a CPU is a totally different ballgame, however.
IMO, if the data can persist on the GPU, then this could be a big win.
And no matter what you do, don't write to same cache lines, especially across NUMA regions. Also avoid locks and even atomic operations. Try to ensure also PCIe DMA happens in local NUMA region.
I'm impressed of getting 100 GBbps CPU bandwidth. It's hard to avoid QPI saturation.
I was involved in a database research project recently, and this is exactly what we found: sure, GPGPUs and the like are much faster than CPUs for the right database queries, but the transfer overhead is so absolutely horrendous that it completely dwarfs any gains in execution time.
cf http://blogs.nvidia.com/blog/2013/11/20/juicing-big-data-sta...
That's not really true; you can kind of treat it that way, depending on the algorithm (eg, matrix multiplication breaks down cleanly that way), but there are serious flexibility advantages in the "single-instruction-multiple-thread" model vs SIMD. For example, consider streaming large numbers of hash lookups - difficult to express clearly with a pure vector processing model.
So: GPUs are wide SIMD machines with a lot of hardware threads, massive branch and glacial memory latencies. When there's a branch or memory latency, HW simply switch thread. GPUs don't care about serial execution performance.
I could get Oracle to sustain 1.2GB/sec read/write pretty easily -- plus four other I/O channels that supported about 600MB/sec each, was fun figuring out the best way to organize everything such that CPU and I/O were optimally saturated.
Throw a GPGPU or Xeon Phi into the mix as well? Fwooaah, fun times.
So asking the folks that already use that stuff could give pretty accurate predictions.
It seems obvious to me that pushing the Foreign Data Wrapper layer with work like this is how we eventually break through the RDBMS scalability barrier of the individual host. In the future, I'm sure you'll see similar work where _the_ GPU won't be the GPU and the PCI bus won't be the PCI bus. Rather, they'll be _a_ host and the network. A database service (database cluster in Posgres's nomenclature) will eventually run not just on a single machine, but on a single cluster of machines. Instead of a cluster of machines for redundancy, you'll have a cluster of clusters.
Postgres is really two things in one: a physical layer of bytes in pages in files, and a logical layer of queries on tables of records. The most important piece in the future will be the logical layer. The FDW layer will naturally be extended and generalized until it is fully as powerful as the current physical layer. At that point, it can be made THE api through which the logical layer accesses data. The current physical layer will then be nothing more than the default implementation of that general API.
At that point, we can move whole or partial tables to other hosts. Perhaps the autovacuum daemon will gain a sibling in the autosharding daemon. The query optimizer will need to care not just about disk IO, but network IO and will need to start considering the non-uniform performance characteristics of different tables. Some tables will be driven by Postgres's default physical storage engine. Others will be driven by other RDBMSs, or by NoSQL key/value or document stores, or other data stores. They may be on the same machine or a different one.
Postgres will transform into a query engine on top of whatever data stores best fit your workload. I expect the query engine will learn about columnar stores and be able to mix those in a single query with the row stores, key/value stores and document stores that it already understands. PostgreSQL will be a central point through which you can aggregate, analyze, and manipulate any and all of your data. It needn't be intrusive or disruptive: you can still use a normal redis client for your redis store, but you can also use Postgres to manipulate that data with SQL and to combine it with other tables, whole RDBMSs, other NoSQL stores, spreadsheets, web services, or anything else. Maybe it will even make things like Map/Reduce frameworks redundant.
I don't typically follow Postgres's internal discussions, so maybe this is already being discussed and planned. Or, maybe it's so obvious that nobody even needs to talk about it. Or, perhaps I'm just some wide-eyed idealist who doesn't understand the fundamental problems preventing such a thing from ever being practical.
But I can't help but wonder what the sys admin's response is going to be when I start asking for additional graphics cards being added to his perfectly built 2U database servers!
http://www.geforce.com/hardware/desktop-gpus/geforce-gtx-780...
I'm not suggesting CPU resources should be counted like that, but that's closer to have GPU resources are counted. Sure, it sounds impressive, but does that 2304 cores really represent fair truth to, say, 4 CPU cores?
NVidia GPUs have a theoretical peak of about 3-5 TFlops for 250 watts. http://en.wikipedia.org/wiki/List_of_Nvidia_graphics_process...
Xeons have a theoretical peak of about 0.5-1 TFlops for 150 watts. http://www.microway.com/hpc-tech-tips/intel-xeon-e5-2600-v3-...
Is that completely apples-to-apples? Probably not since the Xeon is probably talking about double precision floating point versus single precision on the GPU. But for a lot of database applications which don't involve money, single precision floats have a sufficient level of accuracy for the performance improvement to be attractive.
Yeah the performance isn't 100x like it used to be but it's still enough that if you have racks and racks full of machines a 3-10x improvement could be really substantial. Going from $10k/mo in rent to $1k/mo in rent at a datacenter could make or break an early stage startup.
Further as things get cheaper they get used a lot more. Scientists have only two models: the ones they can run but don't really like and the ones they want to run. Adding fidelity to modeling codes isn't an absolute good but it's hard to argue that it makes the world worse.
"2300+ cores" is a VERY misleading way to represent GPU resources. You could also say GTX780 has 12 cores with 1/3rd clock frequency is an equally unfair and equally "true" way to express it, if you were trying to suggest CPUs are "better".
You are also asking the sysadmin to install a closed, unreliable kernel-level piece of software with that GPU.
it pains me to say that fglrx is still years behind them, however.
Unlike all of the closed network equipment they already manage.
Is / will this acceleration be switched on by default?
There's a limit to GPGPU acceleration though. It's the tiny amount of RAM. We need to adopt shared memory architecture like those found in games consoles. A single massive pool of RAM would further unlock potential power.