Lawrence Livermore National Lab's powerful new supercomputer
mercurynews.com
mercurynews.com
https://hpc.llnl.gov/training/tutorials/using-lcs-sierra-sys...
I'm not sure how big the HMC jobs get on these new machines---it depends on the size of the lattice (which gets optimized for physics but also algorithmic speed / sitting in a good spot for the efficiency of the machine).
My understanding is that Spectrum (formerly Platform) LSF was included as part of their proposal.
The biggest difference however, is that the goal for supercomputers such as this is as high of an average usage rate as feasible. The cloud is abysmal for this where it is more for bursting to say 100k cpu core jobs. For systems like this, they'd want the average utilization to be 80%+ all the time. The cost of constant cloud computing like this, even using reserved instances, would be a multiple of the 162 million USD it cost to build this. Also, the IO patterns you'll see for large amounts of data like this (almost certainly many Petabytes) isn't nearly as cost effective as it is to hire a team to build it yourself.
You're right that the IO tends to be very high performance and high throughput, too.
As much as people love to hate them, I'd love to see you get IO profiles remotely similar to what you can get with Lustre or Spectrum Scale (gpfs). They're simply in an entirely different ballpark compared to anything in any public cloud.
LLNL also has national security concerns that are unparalleled by most AWS applications ;)
With the Clos network style topologies that are commonplace in large data centers today, I'm not sure one couldn't achieve decent results in the public cloud.
AWS networking is pretty terrible, but in GCP, I can get 2gbps per core up to 16Gbps for an 8-core instance. For any bare metal deployment, I'm going to be maxed out around 100Gbps which will be close to saturating an x16 PCIe bus.
It's hard to find a dual-cpu frequency optimized processor with less than 8 cores and I'm not sure that'd be cost effective. With hyperthreading, that yields 32 usable cores or around 3.125gbps per core.
Even still, I wager they'd go for better density.
Also, I can get 8 GPUs along with that 8 core/16gbps instance in GCP. Sounds totally doable to me.
AWS looks like a bad deal here.
If you have a variable load, cloud infrastructure may make sense if you can easily auto-scale.
In my experience, most business real world applications are multi tiered applications with variable loads hence are a good fit for cloud infrastructure.
However, attaining the required application flexibility and KPIs for efficient auto scaling is quite hard and require strong functional & technical expertise.
I'm running infrastructure for a SaaS app in k8s. I feel like I'm doing well sustaining >50% efficiency, i.e. all cores running >50% all the time and more than half the memory consumed for things that aren't page cache. Hard to get better efficiency without creating hot spots.
Not a great deal.
Edit: It sure isn't popular or easy to talk about from the looks of things.
http://usqcd-software.github.io/
People use the stack in a variety of different ways---I'll describe my own usage.
There's a message-passing abstraction layer, QMP, sitting over MPI or SMP or what have you (you can compile for your laptop for development purposes, for example). This keeps most of the later layers relatively architecture agnostic.
Over that sits QDP, the data parallel library. Here's where the objects we discuss in quantum field theory are defined. We almost always work on regular lattices. QDP also contains things like "shove everybody one site over in the y direction" (for example).
Finally, there's the physics/application layer, where the physics algorithms live. I am most familiar with chroma. QUDA is the GPU library and can talk to most application-layer libraries and has at least simple solvers for most major discretizations people want to use (it also has fancier solvers such as multigrid methods for some discretizations). Code in chroma by and large looks like physics equations, if you had a pain-in-the-ass pedantic student who didn't understand any abuse of notation.
Chroma can be used as a library, so that for your particular project you can do nonstandard things while leveraging everything it can already do.
Other physics layers include CPS, which grew out of the effort at Columbia with QCDSP/QCDOC, MILC (really optimized code for staggered fermions), and others.
The USQCD stack isn't the only one. Another modern lattice field theory package is grid, developed by a tight collaboration between intel and University of Edinburgh https://github.com/paboyle/grid. There's also openQCD http://luscher.web.cern.ch/luscher/openQCD/
On a POWER8/NVIDIA P100 machine I know QUDA gets 20% of peak, sustained.
I really wish they just said it has XX TB of memory
Maybe I'm grossly underestimating how much data all human written works would actually occupy... but that sounds like the amount of data I could put on a home NAS (frankly I would have guessed my laptop hard drive until reading that comparison).
https://hpc.llnl.gov/hardware/platforms/sierra
Summary:
* 190,080 IBM Power9 CPU cores
* 17,280 NVIDIA V100 (Volta) GPUs
* 125,626 peak TFLOPS (CPUs+GPUs)
Certificates issued by GeoTrust, RapidSSL, Symantec, Thawte, and VeriSign are no longer considered safe because these certificate authorities failed to follow security practices in the past. Error code: MOZILLA_PKIX_ERROR_ADDITIONAL_POLICY_CONSTRAINT_FAILED
Though it's a bit fuzzy. This article is about "Sierra." Department of Energy awarded $325 million to build two supercomputers, "Summit" and "Sierra" both at 150 petaflops each. IBM took $325M, so, it's $162.5M until someone corrects me.
However, the actual proposal it's under (CORAL) included more than that, approaching almost $2 billion in the RFP. They also added $100M for R&D. All these machine systems are considered NRE projects (non-recurring engineering) and contain more than just the hardware and each is estimated to be budgeted at about $400M per system (operational expenses, etc).
Also, the performance of 200K cores on Amazon in VMs compared to 200K physical cores is a lot different. These HPC systems are designed to eek every last 3-5% of performance out of the entire thing, something that you simply can not do even if you try using virtual machines or the cloud.
Every time a new IT manager sees our costs they immediately declare 'I can get you that on the cloud for a fraction of the cost' and starts trying to decommission the system. Never mind the 20-100x increase in solve time and the multi-TB a day data transfer required. Waste 10 hours in meetings, stave it off once again, and gear up for the next round in January...
Modern supercomputers aren't particularly expensive- in the several tens of millions for the capital cost of the machine, several tens of millions for the storage system, several tens of millions for the space, several tens of millions for the power, and a few million for the support staff. In this case, Sierra probably cost $100M for the base machine and storage services.
What is the cutoff?
I've commonly seen 500 microsecond latency between certain AWS zones. That's pretty impressive. Inside a given zone, 200-300 microseconds isn't uncommon.
Will this computer be used close to 100% /was the old super-computer just not enough?
* Weather modeling
* Nuclear Research
* Car crash simulations
* Electronic Design Automation (ex: mathematically proving chips are correct)
* Protein folding: looking for new chemicals for medicine.
Etc. etc.
There's a huge need to grow our supercomputer capacity. I bet you that every major field has a use for a super-computer.
I think there's an element of bragging rights. But the USA buys "practical" supercomputers most of the time. There are designs that push out more FLOPs but are less useful to scientists.
The hugely powerful interconnect and CAPI / NVLink connections on this supercomputer demonstrate how "practical" the device is. Most people are RAM constrained, or message-constrained, and these are the biggest and best interconnects available in 2018.
Interconnects are NOT a "bragging" metric, very few people look at it. Most people look at the Linpack benchmark (purely FLOPs measurement). However, experts can tell when a supercomputer is built with a poor interconnect and purely for "bragging rights" reasons.
How many instances of the Linux Kernel is running in total? Is it only one Kernel instance for the entire machine?
According to https://access.redhat.com/articles/rhel-limits , RHEL 7 on POWER system can "only" manage 32TB of memory, (the whole system has more than a petabyte of memory) assuming they don't run a modified version. So there is definitely not only one OS running, I guess.
Still, the goal of an OS, especially an Unix-like "Time Sharing" OS, is to manage resources. I wonder how hard it would be for the Linux kernel to manage the entire system (with some virtual devices to aggregate all the nodes in one system, even if they are at the opposite side of the room), and if all the developed code about scheduling could be reused at that scale.
By the way, even systems completely across the room may be "close" to each other, depending on the topology of the system. You can imagine this being important for solving a physical system which is periodic in certain dimensions, so there should be little interconnect distance between physically-distant nodes.
It looks more like a private data center with 4000 dedicated machines in the same network to run distributed algorithms than a "single super computer". Are we just "wow-ing" at what is basically a data center here?
If you have an embarrassingly parallel problem, it will run well on this machine but it will also be a waste of the machine's expensive design. Embarrassingly parallel problems run just as well on generic data center hardware. This machine is built for problems that only parallelize effectively with low-latency coordination between nodes. Such problems come up a lot in scientific/engineering simulations but are comparatively rare in general purpose computing environments. General purpose nodes in a cloud computing environment cannot run some of the harder problems this machine runs, at any price. For any non-trivial parallel computing job there comes a crossover point where adding more nodes makes the total time-to-solution longer rather than shorter. This point comes a lot sooner if you don't have dedicated high-bandwidth, low-latency interconnects between nodes.
If they are some info on internet about the software stack/architecture of the entire system, I would document myself on that. I didn't explore all the links I posted above yet.
I'm nowhere an expert, and HPC is really specific use case, but there is surely interesting bits to learn from it