That's a lot of factors to keep in mind when you're spending millions on a supercomputer. I'd go for an established hardware provider as well. I'm not knocking the custom AI chips, but I wouldn't try and max out my budget with those - keep them for smaller applications for now.
What matters is the software platform and compatibility: you don’t want to retool and port all your existing programs. That’s where having a platform - CUDA - becomes a deal-maker.
What I don’t understand is how consumers - big, institutional consumers- don’t insist on viable vendor neutral platforms like OpenCL. Obviously they’re the only ones that have an interest in cross-vendor compatibility.
The same reason a big institutional consumer will happily pay Microsoft for Windows Enterprise, instead of demanding Microsoft rebuild Windows using open source technologies.
A supercomputer or DC cluster has a fixed lifespan and amortization period, if you can get something in a contract, with everything you want for a price you are happy with, why would you introduce risk to it by demanding open standards for no benefit?
Of course you don’t get to dominate a field to such an extent without some redeeming positive quality.
What I’m arguing - or more appropriately, what I’m wondering - is why on Earth aren’t consumer consortiums not insisting on open-standard platforms.
I get how you want to eek every drop of performance out of your tools without making it your only mission (that would be the research you’re using the tool for), but tying yourself to a single proprietary vendor defeats that purpose imho
I also attended a presentation at IPDPS 2018 by someone from one of the big US national labs where they talked about how absolutely huge of an undertaking it was to port their codes to make efficient use of GPUs. You don't just re-do all of that if the payoff isn't enormous.
Is there anything useful someone could do with it or it's just too much of a problem to set it up and repurpose it?
I cannot state with 100% clarity what happens after the systems are acquired by surplus vendors, but I can say that we have received certificates of what was "recycled". Once the surplus vendors take possession of the hardware, it's theirs and they can do what they want with it. In fact, we recently had to purchase EOL'ed, refurbished Infiniband switches to continue to support an interconnect fabric still in use (~2016). Interestingly enough, some of the switches still had "core" and "edge" labels on them.
> Is there anything useful someone could do with it or it's just too much of a problem to set it up and repurpose it?
In my opinion, it really depends on the node. If the chassis supports hot swappable & redundant hardware (HDD, PSU, etc.), then we'll typically cannibalize several chassis to create an administrative and/or infrastructure node. Case in point, a large portion of our older 12 core nodes have been put aside to serve as administrative nodes, all fully redundant (RAID1 HDD's, dual PSU's, ECC memory, etc.). Since we have a stack of these, we're fairly confident that these will serve us for the next few years, worry free, given the abundance of parts lying around.
GPT-2 117M training at 1 million tokens/sec.
Now, I don't have experience with DGX clusters, so I'm not going to make a firm statement. What I will say is that I, as an outsider, managed to achieve a performance level that is ~unheard of for GPT-2 training. And you can too; TPUs are pervasive.
A TPUv2-512 isn't even as far as the gas pedal goes, either. v3-512 can train all of imagenet to 75.9% accuracy in 4 minutes: https://twitter.com/theshawwn/status/1223395022814339073
v3-1024 can do it in 2 minutes: https://twitter.com/theshawwn/status/1234654848114520065
I once attached a debugger to a training run during startup, after the infeed loop began (meaning it was feeding inputs to the TPU, but no training was happening yet; it was "winding up") and was shocked to discover that when I hit c to continue, it trained on all of imagenet in like 54 seconds. That blows the lid off of every perf result here (under "image classification"): https://mlperf.org/training-results-0-6
(It's not a fair comparison, but it was quite astonishing to see the raw horsepower in action.)
So, nVidia has some catching up to do. And I don't know if they'll be able to. The TPU ecosystem may be clunky at the moment, but boy is it effective. Your options are to invest your time in this ecosystem, which will likely be around in ten years, or in DGX-cluster-type knowledge, which ... might be less pervasive in 10 years.
The distinguishing feature of a TPU is that it has a CPU on board. In fact, it has a CPU with 300GB of memory for every 8 cores. Friggin' love these things.
TPUs will probably dominate AI training, but not so much supercomputing.
Graphcore looks great! But no one outside Graphcore has used them so who knows. Intel's Nervana looked great on paper, right up until they dumped it.
There are some interesting accelerator options around for inference. But for training it's NVidia for almost everything, and TPUs as a good option is a few cases. But you can't buy TPUs (except for the inference-only Coral board), and universities like to own hardware.
This part really confuses me. I'd love to try out their hardware, but the only cloud offering they had last time I checked made you rent a whole month's worth (for many thousands of dollars).
It seems like getting it in the hands of devs should be a top priority.