Accelerated Computing Powering World’s Fastest Supercomputer
blogs.nvidia.com
blogs.nvidia.com
There's a new one. Can anyone comment on the utilization of these megacomputers? Do they have somewhere near 100% usage, with a queue extending weeks? Also, is all this computational power really... necessary? I've seen some intensely inefficient simulation code in my time.
[0] http://www.doeleadershipcomputing.org/incite-awards/With those massive clusters you could afford testing in parallel trillions of mutations to your simulation kernels and prove the correctness of the fastest ones -- pruning should be extremely fast by finding counterexamples. Or even higher level architectural optimizations. Surely they could afford at least a few % of the total time on this pre-optimization (although the tools to achieve this automatically would need to be quite sophisticated!).
There are also a few “software synthesis” and “sketching” approaches that use a constraint solver to find all correct implementations of a high level spec, subject to some implementation pattern. Then they either try them all with brute force or pick the one that optimizes some objective function.
Facebook's Tensor Comprehensions framework, which generates CUDA kernels through a genetic algorithm, is closer to the sort of approach which would take greatest advantage of the hardware it's on.
There are separate debug queues for smaller development and test jobs that prioritize latency over throughput, and this is a main source of inefficiency (but, as someone who develops on supercomputers, trust me, it's necessary to make them usable).
As for the efficiency of the codes--there are a lot of barnacle-encrusted Fortran codes in astrophysics, chemistry, and elsewhere. Some of them probably contain plenty of inefficiency, but the cost of re-writing them is high. Oftentimes either the people writing the codes aren't interested in maintaining good software quality (because they're focused on the science) or the incentives aren't aligned to re-write code that's already established, but from the 80s (re-writing old code is generally not how one gets a Ph.D.). But they do very important things, like simulate nuclear fusion or the interaction of chemicals in the atmosphere that create weather patterns and storms, so it's very important to run them.
There isn't necessarily a lot of research into packing these better - the basic algorithms have been unchanged for quite some time, and a lot more effort goes into deciding how to prioritize different groups that are sharing access into the same system.
97% is the highest specific value I can recall for any of the larger sites with a heavily mixed workload, absent having a nearly infinite supply of short+small jobs at hand to use to fill those gaps. And a lot of users aren't trying to chase that - for these large scale "capability" systems the goal is to scale out as large as you can anyways, the smaller stuff is usually relegated to "capacity" systems elsewhere with a less expensive architecture.
One thing that at least some schedulers can manage is the idea of a min+max runtime for a job, combined with a min+max node/cpu count. If you have users willing to 'scavenge' otherwise wasted time by running under such a regime that can put you closer to full usage.
https://www.top500.org/statistics/overtime/
it appears the number of top500 systems with accelerators is not really increasing much - it has been sitting at around 20% of the top 500 machines for the past ~4 years.
Can anyone working with these systems comment on why that might be? Are accelerators still tough to apply to a lot of the problems these machines are used for?
* 250 PB GPFS
* 2.5 TB/s
I'd like to find some information on the burst buffer that is sitting between compute and the capacity storage system.
> Built for the U.S. Department of Energy, this is a machine designed to tackle the grand challenges of our time. It will accelerate the work of the world’s best scientists in high-energy physics, materials discovery, healthcare and more, with the ability to crank out 200 petaflops of computing power to high-precision scientific simulations.
EDIT: OK, it's 200 PFLOPs of high precision math, 3 EFLOPs of lower precision math. I take my comment back.
The 200 petaflops number of Summit is referring to general purpose double precision computation. This is totally incomparable to anything the TPU can do at any performance.
For the doubters and the disbelievers that have been wondering what is the relevance of IBM in this day an age: this. This is what IBM is all about.
And it's not just about PFLOPS; each node has 1/2 terabyte of memory, globally addressable across the entire cluster using RDMA over Mellannox 200Gb/s EDR.
It's also P9: 44 cores per node; but most importantly each node drives a couple of V100 through NVlinks, which allows the GPU to share the system's main memory.
Also, you can defend yourself later when people respond. Prematurely saying “I am ready to take the heat” is not how we communicate here.