2080 RTX performance on Tensorflow with CUDA 10
pugetsystems.com
pugetsystems.com
These numbers match up with the performance that we’ve measured in our own tests that were posted last week. The Titan V is simply too expensive for Deep Learning. The 2080 TI is, by far and away, the best GPU from a price/performance perspective.
As mentioned in the article, only possible reason that you might want a Titan V is if you care about FP64 performance: i.e., nobody training neural networks.
Forms matter. Colloquial meanings also matter, but not as much, particularly when they're an egregious violation of English and decency.
"When I first looked at fp16 Inception3 was the largest model I could train. Inception4 blew up until I went back to fp32. Mixed precision needs extra care, scaling of gradients and such. Still I think it is a good thing. What I really want to test is model size reduction for inference with TensorRT targeted to tensorcores. I think that is probably the best use case. Non-linear optimization is just too susceptible to precision loss."
There was also some NVidia video presentation recommending mixed FP32/FP16 training instead of pure FP16.
I am doing experimental work where I really need to have double precision i.e. FP64. The Titan V offers the same stellar FP64 performance as the server oriented Tesla V100.
Still, I think 2x1080Ti is a better deal than 1x2080Ti and costs the same.
Also the 2080ti can do lower precision math (int8/4) in the tensor cores, while the Titan v cannot.
I'm using my GPUs to train large sequence to sequence models (with long sequences) that need FP32 for training and can use FP16 for inference (mixed-precision training), so I can't even use the FP16 performance of the Tensorcores for training.
The only disadvantage is that the energy costs are higher using two 1080 Ti's compared to one 2080 Ti.
Example: https://www.pugetsystems.com/nav/peak/tower_single/customize...
[1] https://www.intel.com/content/www/us/en/processors/xeon/xeon...
sgemm on 5000x5000 matrices takes about 600ms on a Threadrippers 1950x, but only around 150ms on the comparatively priced i9 7900x. Vector libraries for special functions, eg Intel VML or SLEEF also provide a similar performance advantage there.
If you're mostly crunching numbers, and either compiling the code you run with avx512 enabled (eg, -mprefer-vector-width=512 on gcc, otherwise it's disabled) or using explicitly vectorized libraries, you will see dramatically better performance from avx512, regardless of any thermal throttling. Number crunching is what it's made for.
Granted, you should be offloading most of those computations to the GPU, which will be many times faster. But I'd you're in the business of ML or statistics, I'd still way that more heavily than the difference in how long it takes them to compile code.
I don't follow the logic. It sounds like you're saying that if you care about that specific type of highly vectorized computation being fast what you really want is a GPU rather than any particular CPU. So how should that have a major influence on which CPU you choose? Particularly when the CPU which is slower at that is faster at many other things that aren't suitable for a GPU.
If your number crunching is just neural networks on your GPU, then the CPU doesn't matter.
But there's probably a lot of overlap between the folks who train neural networks, and those who may do linear algebra, MCMC, or traditional stats that are much better suited to the CPU. That is, conditioning on person A being someone who trains NNs, there is a higher probability that they're someone who would be interested in CPU intensive tasks that benefit from vectorization. If that isn't you, don't factor it into your decision.
I do most of my number crunching on the CPU, so my choice is clear. The reviews of avx512 are generally poor (disable it so you don't get thermal throttling!), while the Threadrippers receive a lot of praise. But within it's own niche (linear algebra, many iterative algorithms), the widest vectors are king.
I think you're also looking at the release prices for the CPUs rather than the current ones. Using today's prices from Newegg, the Threadripper 1950X is $699, the (newer/faster) 2950X is $859, meanwhile the i9-7900X is $1275, up from its $989 release price presumably due to Intel's current manufacturing issues. And the AMD processors have 60% more cores/threads with, avx notwithstanding, generally equivalent performance per thread.
I expect you're right that there are niche workloads where avx512 is a real advantage, but it's starting from a pretty deep hole on the price/performance front in general.
With FP16 one can theoretically get 2x speed and 2x larger models with the same VRAM capacity. For inferencing with INT8/INT4 it can be even way better (good for embedded stuff). The downside is that sometimes more complex/deep models don't converge (or converge less often than FP32). Sometimes there are framework issues with some advanced FP16 stuff.