High performance models in TensorFlow
tensorflow.org
tensorflow.org
Is there a motherboard out there that would allow the usage of 8 GPUs with each of them allocated the full 16 lanes?
The best I've found so far would be 4 GPUs on one board, either with 16/8/8/8 or in a rare case 16/16/16/16 (but that requires 2 CPUs).
Besides the physical space, which again seems to be limited to 4 double-wide GPUs on one motherboard.
Another really nice one is this one:
https://www.supermicro.nl/Aplus/system/Tower/4021/AS-4021GA-... (AMD based)
https://www.supermicro.nl/products/system/4U/7047/SYS-7047GR... (Intel based)
http://www.dell.com/us/business/p/poweredge-t630/pd
(That last one should be able to hold 4 GPUs but I'm not quite sure about whether or not it will be able to power all of them, Dell isn't helping with their documentation either.)
Edit: just found this:
http://www.supermicro.nl/products/system/4U/4027/SYS-4027GR-...
About $4K, + 8 GPUs that's a pretty penny. Drool.
I was a little worried because this is the most power-hungry machine I've ever built, but so far I haven't had any issues. But I'm not going out of my way to start an electrical fire, either.
The motherboard in question is pretty deluxe, the only thing I'd complain about is the boot time.
Z10PE-D8 WS mobo, 2 x Xeon E5-2620 v4 (32 threads), 128 GiB DDR4 ECC RAM (8*16 GiB), 256 GiB SM961 NVMe, 12 TiB SATA (4 x 3 TiB HGST server drives), 3 x GTX Titan 6GB, Soldam black knight XR1 casing, 2 x DeepCool Gammax 400 CPU coolers, EVGA SuperNOVA 1300 G2 PSU, 2 x Noctua NF-A9 PWM, 2 x Noctua NF-A14 3000 PWM
A GPU has it's own processor and RAM. If you're transferring to your system RAM and back again often enough to max out x16 PCIe lanes you should fix that.
If found a tool to monitor this unfortunately it only works on Xeons and not on what's in my desktop.
It also depends on what you plan to do with the gpu. For example, models that do most of the work on the gpu and rarely ingest data from the host, such as large and slow models, will run just fine. On the other hand, attempting to parallelize training across GPUs and nodes is a chore...
There are quite a few examples of folks keeping 8 GPUs busy. Eg. Baidu's speech recognition training (which uses quasi-rnn IIRC).
Using external GPUs over USB-C gives useful speed for training NNs.
Above that, you need to look at c612: https://www.supermicro.com/products/motherboard/Xeon/C600/X1...
But I agree with your assesment, I have noticed several barely interesting blogposts/arxiv papers upvoted.
A more extreme version of this is situation is of the medical researcher who "reinvented" on his own the trapezoidal rule for numerical integration[1].
[0]: https://xkcd.com/1053/ [1]: https://fliptomato.wordpress.com/2007/03/19/medical-research...
Looking at [2] without having context of [1] can be confusing. But no one is trying to pass this off as "innovation".
[1] https://www.tensorflow.org/performance/benchmarks [2] https://www.tensorflow.org/performance/performance_models