Nvidia Announces PCIe Tesla V100
anandtech.com
anandtech.com
The upcoming Volta-based consumer GPUs are going to be your best value for machine learning, not V100.
[1] https://www.amazon.com/NVIDIA-Tesla-P100-computing-processor... [2] http://accessories.us.dell.com/sna/productdetail.aspx?c=us&l...
http://images.nvidia.com/content/pdf/quadro/data-sheets/3020...
Agreed re consumer GPUs for machine learning though.
I'm not sure how consumer GPUs that don't exist can be better, though. That seems highly presumptive, especially given that most machine learning setups are heat, cooling and space limited.
You forgot "cost". V100, like P100, will cost literally more than 10 times what consumer cards do, for a likely speedup of less than 2 times. For the vast majority of people, that will not be worth it.
Of course eventually maybe they'll come out with a cheap but powerful deep learning GPU (though with the V100 having a 800mm2 die, inexpensive is going to be relative), and it's impossible for me to compare with future products, but there's a good reason these cost as much as they do.
The gap between V100 and 1180 is yet to be determined and depends mostly on what Nvidia does with the tensor units. We shall see. But I am extremely confident in saying that the performance gap will be nowhere near 60x, let alone higher than that. And despite the early announcement of V100, 1180 is likely to be available to most people long before V100, just like P100 and 1080 before.
I said that FP16 performance is at least 60x worse, which it absolutely is. The cores on the GT102 do not natively support FP16, so each FP16 operand has to be converted to FP32 and then processed, with significant overhead. The P100 can use FP16 directly yielding not only double the processing speed (because the FP32 cores are really dual FP16 cores, like registers on an x86), but a major memory savings. The 1080(Ti) is crippled at FP64 as well, again dramatically slower than the P100. It's a consumer GPU and FP16 and FP64 just aren't usually a consumer need.
These aren't just the same thing, with one having a sucker price. They are dramatically different chips, and many applications require the P100.
just like P100 and 1080 before.
Comparing the P100 and 1080Ti as if the latter is the consumer version of the former is not useful. They are profoundly different chips.
OK, I didn't think that's what you meant because it's a rather silly benchmark to use as it artificially disadvantages the 1080. FP32 deep learning performance on 1080 Ti will usually be ~half of P100's FP16 performance. FP16 is an advantage, but not a 60x advantage, and the 15+ 1080 Tis that you can buy for the price of one P100 are going to be far faster in most deep learning scenarios (admittedly at a higher power/etc cost).
> Comparing the P100 and 1080Ti as if the latter is the consumer version of the former is not useful. They are profoundly different chips.
I couldn't disagree more. The chips are similar enough that most deep learning applications could run on either. Really the only reason to choose P100 for deep learning with the ridiculous prices Nvidia is charging would be memory capacity and bandwidth with the FP16 advantage.
> I couldn't disagree more. The chips are similar enough that most deep learning applications could run on either. Really the only reason to choose P100 for deep learning with the ridiculous prices Nvidia is charging would be memory capacity and bandwidth with the FP16 advantage.
Sure, but there are more applications than just deep learning which is where things get fuzzy. The P100 and 1080TI are fundamentally different for anything relating to FP64. I think the point of the GP is, or would hope it is, not to collapse comparisons to the case of deep learning when things like scientific computing make the comparisons necessarily more nuanced.
Is that double precision?
The Nvidia 1080 Ti has a double precision performance of 332 GFLOPS [1]. If the above number is for double precision computing, the Tesla V100 (PCIe) would be about 337 times as fast (!!)
Does anyone have more insight into these numbers?
[1] https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_proces...
UPDATE : I should have read the article more carefully, it seems to be a mix of FP16 (half precision) and FP32 (single precision). That would likely mean a factor ~10 in computation performance (specifically for deep learning)
Source of this data is my own calculations using historical data on mining difficulty for the various coins, plus benchmarks for typical/optimal-ish GPUs used to mine each coin.
It's possible, though, as you suggest, that there are some folks running ASICs or improved hashing algorithms at a small scale, small enough not to overwhelm profitability of GPU mining but large enough to muddy calculations which assume that all mining is being done on GPUs.
But for many other coins, there are two big factors bringing GPUs back into the game for mining:
- Lots of alt-coins competing for popularity; many might have ASIC mining implementations in the future, but there either hasn't been enough time or enough money to be made by doing it yet.
- There's been various research done to reduce the potential gain in efficiency from an ASIC implementation over using commodity hardware. This was done in response to the concentrating effect of ASIC mining operations, which put a greater percentage of global hashpower in a smaller set of hands. The main way ethereum implements this "ASIC resistance" is by exercising memory bandwidth -- an area where GPUs are already quite optimized -- in the hashing algorithm. https://github.com/ethereum/wiki/wiki/Ethash goes into detail.
A Raspberry Pi has way more compute power than the first generation Cray machine and a 386-vintage machine has better single-core, non-vectorized performance.
Would something like that be even possible with GPUs? I know the Exascale paper from AMD mentions GPU chiplets which sound like this.
I think the Tesla cards in particular are double precision, which is not necessarily the best for deep learning applications.
Here is a starting point: https://en.wikipedia.org/wiki/Multiply%E2%80%93accumulate_op...