DeepLearning11: 10x Nvidia GTX 1080 Ti Single Root Deep Learning Server
servethehome.com
servethehome.com
And that's because to the best of my knowledge, most of the CUDA ecosystem out there was developed on GeForce GPUs.
There is currently no GeForce equivalent of Volta at a time when the underlying programming model has undergone some traumatic changes that really alter the way to write efficient code going forward. If the only way to access Volta GPUs turns out to be AWS instances at $25/hour or $150,000 DGX-1V servers plus hosting costs, I suspect a lot of existing CUDA code will bitrot.
Imagine a near future where AMD Vega GPUs are faster than GTX 1080 TI at FP16 training and inference for deep learning. Without some sort of successor to that GPU, I really think that could happen because Nvidia went out of its way to cripple FP16 performance on GeForce.
That said, at Vega's 26 or so FP16 TFLOPS, it wouldn't be hard for NVIDIA to release a GeForce Volta with 30-40 tensor core TFLOPS, that both stomped on Vega and remained significantly inferior to V100. Given how hard it is to program Volta optimally, I'm surprised they haven't done so already, if only as a Titan XV Edition that can only be purchased from their website.
BTW - do you happen to know what DL frameworks currently support mixed precision training with Volta tensor cores? Curious to see if AWS V100 instances can really do 120 TFLOPS as advertised. I think latest versions of CUDA/CuDNN support Volta now?
https://www.pcgamesn.com/nvidia-geforce-server
And while there is no GeForce equivalent to Volta today, that will not be true in the near future. At some point they will come out with a new GeForce line of cards. In the past, the GeForce generation was either before the Teslas, or just slightly after. I also don't agree that there is nothing competing with the voltage right now on the GeForce line. The 1080ti is not the same performance, but if you are willing to have multiple cards, two of them are just as good or better than the V100.
There's https://rocm.github.io/index.html but it's not quite clear to me how far they got and how usable it is today.
> NVIDIA specifically requests that server OEMs not use their GTX cards in servers. Of course, this simply means resellers install the cards before delivering them to customers.
The 1080Ti card is about 60% of the performance of the P100, but costs 700 dollars instead of 5k dollars. Of course, people will try and build these boxes, they have a much higher ROI compared to Nvidia's DGX-1. So what do NVidia do? Try and stop vendors from selling them! https://www.pcgamesn.com/nvidia-geforce-server
There are plenty of things to arbitrarily segment in a GPU - HBM, large memory sizes, tensor cores, FP64.
The problem with hard disk segmentation wasn't the idea of segmenting at all, it's that there were no good ways of doing it. A higher MTBF? You're better off buying more, cheaper disks. Density (eg from helium)? Might be worth a 20-50% premium in $/gb, but not 200-500% $/transistor as with Xeon.
With disks there was always just as much pressure on the consumer side to keep energy down, cost down, and capacity up, which means that there was no natural segmentation, and no straightforward unnatural segmentation.
I'd like to ah, learn about machine learning so here's to hoping Nvidia doesn't nerf its deep learning capabilities on the driver level.
TLDR; You can scale out distributed Tensorflow training to tens/hundreds of GPUs with AllReduce on machines like this one, not just on the DGX-1.
Are there no appropriate mainboards yet?