Rosetta: The Engine Behind Cray’s Slingshot Exascale-Era Interconnect
fuse.wikichip.org
fuse.wikichip.org
I suspect the answer is in the networking, what this article spends most of its time talking about. HPC networking is so ridiculously much better than commodity networking. 1us latencies vs 100us+, 200Gbps bandwidth vs 10Gbps. I imagine certain kinds of especially simulations involve a ton of communication between nodes simulating nearby cells and everything is completely bottlenecked on that. It might even be that you can't even match a supercomputer's speed by just buying more AWS instances because the scaling factors on the communication mean more instances doesn't make your sim faster. Again this is mostly inference on my part from things I've read, no first-hand knowledge.
I've even seen a few places that have even purchased the prerequisite hardware but then never turned the features on.
Things like hardware flow control, cut-through switching, jumbo frames, SR-IOV, RDMA, etc... are just a few checkboxes away and can give even small VMs crazy good bandwidth and latency.
I've seen jaw-dropping performance when it was set up correctly, something like 99% of the theoretical wire rate of 20 Gbps. With dual NICs and Windows 2016 we got nearly 40 Gbps because it can use both NICs at once! Amazing file copy times. I tested it with the usual 4GB ISO file copy test and it was so fast there wasn't even a progress bar!
Meanwhile, on most networks I see about 1.5-3.0 Gbps effective across dual 10 Gbps NICs. If you think about it, that's crazy bad. It means that in effect they're getting less than 10% of the available physical capacity for data traffic. The other 90% is being thrown away by "passive" secondary cables or by the lack of configuration on the switching layer.
And regarding that 1.5-3.0Gbps across 2x10Gbps... a) wat?!?! and b) I wonder if the number of horribly misconfigured networks exactly like that are out there giving 10Gig a bad name and hurting adoption.
Unless you are in a very controlled environment it is VERY easy to actually slow down a network while trying to make it more efficient.
Still, if folks can deal with those short distances there's lots of inexpensive equipment out there.
Not going to get that sort of latency without extreme custom chips though.
I generally don't get to play with hardware much these days, but my "dream" on-premises kit would be a hyperconverged cluster with AMD EPYC 2 CPUs, NVMe storage, and 400 Gbps networking. It would be interesting to see what kind of performance that would enable compared to the typical cloud platforms everyone is so excited about...
The majority of HPC systems probably use fat tree topologies (including Sierra and Summit as far as I remember). Something to hand which compares fat tree, hypercube, and the dragonfly topology discussed in the article is http://www.hpcadvisorycouncil.com/events/2015/swiss-workshop...
Yes, because that's where you get innovation in architecture. You have to understand the reason commodity hardware caught up supercomputers was because of getting lucky with materials. AKA the mateirals cpu's were made out of was easily shrinkable and frequency kept going up because of dennard scaling until roughly 2006. With the end of dennard scaling, research into new processor materials and computer architectures is the main issue. Single threaded performance is a huge priority because most problems we are interested in cannot be made parallel and hence do not benefit much from multicore.
There are also things you are not hearing about and aren't privy too.
Basically computers can run much faster if you use incredibly expensive or experimental materials. The major roadblock is really mass production. I'm certain there are black projects doing all sorts of custom hardware behind the scenes, most likely for military applications.
But the expense to run those custom machines makes it non viable for anything like mass adoption.
No. The CPUs on supercomputers are basically the same as what you can purchase commercially. What distinguishes supercomputers from "the cloud" is better networking.
> Single threaded performance is a huge priority because most problems we are interested in cannot be made parallel and hence do not benefit much from multicore.
The few big problems very interesting to militaries are all embarassingly parallel. There are no super-secret CPUs that have better single-threaded performance than what is available commercially.
For example, the Rapid Single Flux Quantum (RSFQ) logic uses superconducting circuitry and experimental chips have hit 1 THz for simple ADC/DAC circuits, and I've heard of 100 GHz DSP chips in practical use. They're typically used for radio telescope amplifiers and digitisers, large static military radar installations, etc...
People sometimes forget that the commodity market seeks a kind of Pareto Frontier optimisation where "cost effectiveness" and "low risk manufacturing" are key metrics. If you don't care about any of that and just want the maximum performance at any cost, there's some amazing technologies out there!
That being said, I'm also on your side in guessing that the problems that are money-wise worth running on a supercomputer are getting pretty limited as commodity hardware improves.
As others have said, it is the interconnect that makes HPC special. The ability to run a job across tens, hundreds, thousands of nodes, where all nodes need to exchange large amounts of data with all other nodes every millisecond, without network being a huge bottleneck that leaves your CPUs idling.
If your job is never big enough to outgrow the biggest and baddest single server money can buy, then indeed you don't need HPC. But if you are running e.g. a full fluid dynamics simulation of an F1 racecar, a wind turbine, or a jet airplane, you sure do.
Clearly the customers (mostly governments) think it's worth it. They also buy commodity clusters in addition to supercomputers so they're very aware of the differences. The biggest difference is the network and that's why Cray is developing that part themselves; some simulations are network-bound and using a commodity network it would either cost more than a supercomputer (due to lower efficiency) or it would never reach the desired performance at all (due to poor scaling).
If you have hundreds of 80.000$ nodes with 8 V100 each consuming thousands of Watts per node, you start thinking about how much money is the 90% idle time of those nodes, waiting for data from the network, is costing you (probably millions of dollars per year).
So yeah, a network that's 10x faster is worth every penny if it can turn the system from a 90% idle to <10% idle.
That's why people always pay extra for NVLink, HDR Infiniband or HPC Ethernet interconnects, and also why Nvidia bought Mellanox last month...
Essentially, when your machine has 1 ExaFLOP/s of compute performance, every second that it remains idle waiting for the network you... well lose 1 ExaFLOP. That's just a lot.
Also, imagine millions of cores writing to the same harddrive simultaneously - hell try doing `ls` on a normal linux machine on a folder with 10 million empty files. The "storage" infrastructure to support instantaneous I/O on these systems is quite complex as well. Every second that the system is writing to disk you are wasting another ExaFLOP of compute.
Many applications running on a Top 10 system of the Top 500 are required to show good single threaded performance and good weak and strong scalings up to a big part of the system.
Checking all these 3 check boxes is hard to do if your communication is not fully overlapped by computation. The smaller your problem, the better your single threaded performance (e.g. at some point your problem fits in L3), and the larger fraction of execution time that communication dominates.
Even the poorest applications on these systems use non-blocking MPI to overlap communication with computation nowadays.
If you are building a system whose purpose is to transfer data (e.g. a router), every second that system is "computing" and not "transfering" you are paying for Watts that aren't doing anything useful.
A supercomputer purpose is to compute stuff, not to move data around, so making that argument about a supercomputer doesn't make much sense because nobody cares whether the supercomputer is using all its bandwidth or not as long as its using all of its compute capabilities.
So that's like making the argument that everything that isn't a plane, well, isn't a plane. Sure, my guitar isn't a plane, but it is not the purpose of my guitar to be one :D
Two Xeon 8280 CPUs (28 cores each), 10 nVidia Tesla V100 32GB cards 1536GB DDR4-2933 memory, registered ECC, 24 DIMMs 4x 3.8TB SSD, SATA, Intel D3-S4610 Series and a Mellanox ConnectX-5 IB InfiniBand, dual port EDR 100Gb/s
is ~$160K.
That is indeed some very expensive idle time (although, I think I'd work as hard as possible to get my problem to fit on individual machines with data on external fast storage).
FWIW I think it would be rare for somebody to build that kind of node. Applications that fully utilize two CPU sockets while simultaneously maxing 10x V100s are super rare. Usually for GPGPU applications the CPU is mostly idle, and only "orchestrates" the GPU usage.
If one needs to support both GPGPU-only and CPU-only applications in a cluster, one would usually just build two separate partitions. One with e.g. 4-8 V100s per node, but way weaker and cheaper CPUs, and the other one with 2 beefy CPU sockets but no GPUs.
Also, if you are going for that kind of interconnects, you are probably building a system with >100 nodes. For your price, that's 160 million dollars, which is on the ballpark of what these systems costs. You usually get a cheaper price per node at that scale (although not much cheaper), but there are many costs (installation, maintenance, training, etc.) that are not covered there. In particular, those nodes will be running quite hot (just add the watts that those V100s and the CPUs use!), so the cooling infrastructure required to held those in a relatively small space can also end up costing quite a bit. You don't want a water pipe breaking and destroying a rack containing 10 of these nodes costing 1.6 million dollars, right?
As for your comments on water pipes breaking: I work on Google's TPUv3, which uses water cooling to the TPU chips, see https://www.nextplatform.com/2018/05/10/tearing-apart-google... for some speculation on the details.
This depends on both the application and the actual machine. Moving data from storage or from the network to the CPU to then feed it to the GPU just increases latency. Unless you are doing some non-parallelizable pre-processing of the data in the CPU, you are often better off with doing DMA or RDMA directly to the GPU memory, fully bypassing the CPU.
Not all machines support that, and not all applications support that either, but many applications do. Arguably, even for the applications that do this, leaving the CPUs unused is not really a desirable property. I've seen a couple of applications that split the compute workload to the CPU to fully utilize the CPU as well, and I've also seem other applications that once the GPU finishes processing they move the data to the CPU memory to leave space on the GPU for the next batch as quickly as possible, and the CPU then streams the data off the node while the GPU continues to compute (so you get an RDMA->GPU->CPU->Storage/Network loop).
> see https://www.nextplatform.com/2018/05/10/tearing-apart-google.... for some speculation on the details.
Thanks for the link, I never learned much about this part of the system.
There just aren't that many codes that can take full advantage of these monsters, and there are other, cheaper, faster ways to obtain roughly the same results. For example, in my ex-field of molecular dynamics, it seems like you can get away with minimal interconnect by running many simulations and pooling statistics across runs, which needs a bunch of IO bandwidth but doesn't need supercomputer-style interconnect.
What's interesting is that supercomputer hardware is really great for high performance ML training, but the supercomputer folks came to that realization later, and only after companies like Uber ported the tensorflow distributed implementation to MPI (Horovod). And most of the ML people were poking around with slow interconnects. Now it's converging- the largest use of supercomputer-style systems in 10 years will be ML training, not physical simulations.
Supercomputer spend has only limited impact on forward technological process.
Which beats using a physical model and highspeed film projected onto a wall covered in graph paper and manually digitising data.
My bitter experience, particularly with "big data" people, is that they just won't be told by those with long relevant research computing experience in HPC and similar. It's hardly that HPC people don't know what systems can do, particularly if they live and breathe things like distributed linear algebra. I despaired, and largely gave up, when someone strode into a chair From Industry and assured us that the university didn't do Big Data, notwithstanding LHC, astronomy, sequencing, synchrotron work, etc., and was then going to build a Big Data Cluster from a few knackered PCs (and in a basement!) to better the HPC systems. Then I had MPI explained to me.
so really a port can fall back to supporting ethernet? maybe it would have been easier to just put an adapter in there?
just wondering what the real meat is. i do think convergence with non-supercomputer systems is a great idea and should help quite a bit with NRE