World's fastest radix sort? 1B 32bit keys a second using a stock GTX 480
code.google.com
code.google.com
Disclaimer: I didn't follow his work closely, so my interpretation may be misleading.
Some people are willing to pay a premium for the above 3 advantages, that's why Nvidia can sell Tesla cards at such a high price. IMO, the first advantage is overrated: Nvidia's ECC implementation (secdec) is unable to detect many multi-bit errors. Gaming cards offer at 10x better perf/price ratio, so it makes it hard to ignore them.
The Quadra has ECC, it's produced by NVIDIA itself (not 3rd party), and they have developed extensive driver support where each application that you run (Windows) would have specific settings that give you the best performace, quality or stability.
There are tons of other manifacturers that produce NVIDIA cards, but they don't put ECC memory, boost the GPU frequency until they almost don't burn, etc.
Also some specific features exist only for Quadra, that are not important for games - for example quick antialiased lines (just a simple example, there is more).
So what MRB is saying is right. At the surface it looks Quadra/Tesla are worse, but their stability is much better, and customer support is guaranteed.
NVIDIA can't simply support HW not built by them in large. It's not possible for them to know every kind of card there is.
That said the 20- series is really freaking expensive. When I got my C1060's, they were ~$1500/per retail; I think the C2070's are nearly twice that.
That might be hard when going for reliability. It's a bit like overclocking your CPU, you can do it, but it's outside of the warranty. In this case the overclockers are the third party board manufacturers that are trying to get an advantage over their competitors that they can paste on the box so that in the store when comparing two boxes the customer will pick theirs.
> (Also, the value of ECC is diminished if a Tesla is over 3x the price of a GeForce, because then you can run three times and vote.)
That's not really true. The majority of the situations where the Tesla boards are used you'll find multiple boards in one box already. The limiting factor here is how many boards you can pack in to a system, more = better. So if you'd use the 'voting' principle you'd be wasting tons of power, rackspace and host computers in order to offset the price of the ECC, not just the two extra cards.
I don't think so; most GeForce cards appear to use stock clocks and Nvidia reference PCBs and coolers.
I expect we'll see the exact same thing happen as what happened with the 200 series, the first cards will be looked at as the crummy old ones as soon as the next batch hits the streets. You can pretty much expect the following:
a 1.4 GHz part
maybe a 1.8 or 2.0 GHz part
a dual chip card featuring 896 cores at the same power consumption level as the current 480GTX
Edit: Wikipedia says that the Tesla cards have 4X higher double precision performance than the game cards, so that could be the explanation.
Of course there is overhead, that's obvious, there would be overhead in any co-processor driven situation, but that overhead depends to a large extent on the host machine and the bus used to connect, so it would be reducing the value of the benchmark to include those figures in the timing.
Also, there are plenty of applications where the input to the radix sort would come from other kernels and/or where the output would go to other kernels.
In those cases there is no overhead.
Sorting a billion of anything in 1s seems likely to kick against other problems.
Just curious (I'm not asking why you'd want to, but you can tell me that too if you like).
[edit] I don't know how practical that is on real world hardware; a billion keys is obviously a lot of data to transfer.
e.g., consider z-sorting vertices for a million-polygon model as it rotates (that's useful for some kinds of transparency.) Bandwidth to main memory is a small piece of the total.
However, a more recent paper is Accelerating SQL Database Operations on a GPU with CUDA: http://www.cs.virginia.edu/~skadron/Papers/bakkum_sqlite_gpg...
The references in that paper are pretty interesting, too.
SQLite is designed for embedding into cell phones and web browsers and other tiny devices, so I'd expect that there's quite a bit of room for optimizing the floating point math on server-grade hardware.