The Frontier computer, which broke the exascale barrier in 2022
bloomberg.com
bloomberg.com
https://www.tomshardware.com/news/two-chinese-exascale-super...
> Frontier uses 9,472 AMD Epyc 7A53s "Trento" 64 core 2 GHz CPUs (606,208 cores) and 37,888 Radeon Instinct MI250X GPUs (8,335,360 cores).
https://www2.cisl.ucar.edu/about https://www2.cisl.ucar.edu/ncar-supercomputing-history/bluef...
I don't know if the top ones are still used for oil and gas exploration (crunching data to provide higher resolution and higher accuracy oil field maps), but they have been in the past.
Less politely, it exists because of The Bombs. Everything else it will do is a hobby. It is specifically designed for whatever compute needs their classified program has.
https://en.wikipedia.org/wiki/Stockpile_stewardship
https://www.energy.gov/nnsa/articles/stockpile-stewardship-a...
All of this capital is being spent because the test ban treaty prohibits the detonation of these weapons, but the US stockpile and its efficacy must be maintained and new weapons must be designed. So how do you know if an old weapon or a new design will work or not?
Some part of the answer is real data + simulations. Lots and lots of simulations. It's why Frontier exists and why the DoE has pursued exascale compute for quite some time.
(as for the real data part, it's why places like the National Ignition Factory exist, and why there's an active experimental program that exists to study dummy warheads in unexpected ways, https://www.youtube.com/watch?v=FYdAT0v4DHs )
———
Addendum, re: comments arguing that the NNSA computers are separate etc., the points raised are both true in the specific sense, but not true holistically.
Quotes from the most recent report to congress, published in March 2022 about the Stockpile Stewardship and Management Plan,
Near-Term and Out-Year Mission Goals:
◼ Advance the innovative experimental platforms, diagnostic equipment, and computational capabilities necessary to ensure stockpile safety, security, reliability, and effectiveness:
– Achieve exascale computing by delivering an exascale-capable machine and modernizing the nuclear weapons code base
– Develop an operational enhanced capability (advanced radiography and reactivity measurements) for subcritical experiments
– Quantify the effects of plutonium aging on weapon performance over time
– Assure an enduring, trusted supply of strategic radiation-hardened microsystems
and, The weapons comprising the U.S. nuclear stockpile are assessed to be safe, reliable, effective, and secure. DOE/NNSA’s scientific infrastructure is currently adequate to support stockpile actions. The DOE/NNSA plans to address near-term gaps in required scientific capabilities by deploying the DOE/NNSA’s first exascale computing platform in fiscal year (FY)2023 and by improving capabilities to conduct subcritical experiments through the Enhanced Capabilities for Subcritical Experiments project by FY 2026.
It seems they're building three of these computers, Frontier is one of them, from the budget requests and a press release, The budget request for Advanced Simulation and Computing increased to support pursuing new validated integrated design codes and advanced high-performance computing capabilities, including the El Capitan exascale system procurement.
The press release, Featuring advanced capabilities for modeling, simulation and artificial intelligence (AI), based on Cray’s new Shasta architecture, El Capitan is projected to run national nuclear security applications at more than 50 times the speed of LLNL’s Sequoia system. Depending on the application, El Capitan will run roughly 10 times faster on average than LLNL’s Sierra system, currently the world’s second most powerful supercomputer at 125 petaflops of peak performance. Projected to be at least four times more energy efficient than Sierra, El Capitan is expected to go into production by late 2023, servicing the needs of NNSA’s Tri-Laboratory community: Lawrence Livermore National Laboratory, Los Alamos National Laboratory and Sandia National Laboratories
El Capitan will be DOE’s third exascale-class supercomputer, following Argonne National Laboratory’s "Aurora" and Oak Ridge National Laboratory’s "Frontier" system. All three DOE exascale supercomputers will be built by Cray utilizing their Shasta architecture, Slingshot interconnect and new software platform.
Link to the congressional report, https://www.energy.gov/sites/default/files/2022-03/FY%202022...Link to the press release, https://www.llnl.gov/news/doennsa-lab-announce-partnership-c...
-
It does seem to me that it's an accurate assessment to say that this project exists because of the nuclear weapons research mandate.
They're the same computer design, replicated. Are they using these to test/mature their software architecture? Hoping that many eyes will make bugs shallow? Or, is it something else?
Perhaps it's seeing patterns where there are none, but for me the link between these projects and stockpile stewardship seems to be undeniable.
DoE is building (and has built) similar leadership-class computers for stockpile stewardship but it's a completely separate program that actually usually lags behind the open science machines they build. El Capitan is the system they're building at Livermore for stockpile work.
I don't believe that Frontier was used directly for COVID vaccine research, but it very well could have done.
Here's the most recent report to congress, published in March 2022 about the Stockpile Stewardship and Management Plan,
Near-Term and Out-Year Mission Goals:
◼ Advance the innovative experimental platforms, diagnostic equipment, and computational capabilities necessary to ensure stockpile safety, security, reliability, and effectiveness:
– Achieve exascale computing by delivering an exascale-capable machine and modernizing the nuclear weapons code base
– Develop an operational enhanced capability (advanced radiography and reactivity measurements) for subcritical experiments
– Quantify the effects of plutonium aging on weapon performance over time
– Assure an enduring, trusted supply of strategic radiation-hardened microsystems
and, The weapons comprising the U.S. nuclear stockpile are assessed to be safe, reliable, effective, and secure. DOE/NNSA’s scientific infrastructure is currently adequate to support stockpile actions. The DOE/NNSA plans to address near-term gaps in required scientific capabilities by deploying the DOE/NNSA’s first exascale computing platform in fiscal year (FY)2023 and by improving capabilities to conduct subcritical experiments through the Enhanced Capabilities for Subcritical Experiments project by FY 2026.
It seems they're building three of these computers, Frontier is one of them, from the budget requests and a press release, The budget request for Advanced Simulation and Computing increased to support pursuing new validated integrated design codes and advanced high-performance computing capabilities, including the El Capitan exascale system procurement.
The press release, Featuring advanced capabilities for modeling, simulation and artificial intelligence (AI), based on Cray’s new Shasta architecture, El Capitan is projected to run national nuclear security applications at more than 50 times the speed of LLNL’s Sequoia system. Depending on the application, El Capitan will run roughly 10 times faster on average than LLNL’s Sierra system, currently the world’s second most powerful supercomputer at 125 petaflops of peak performance. Projected to be at least four times more energy efficient than Sierra, El Capitan is expected to go into production by late 2023, servicing the needs of NNSA’s Tri-Laboratory community: Lawrence Livermore National Laboratory, Los Alamos National Laboratory and Sandia National Laboratories
El Capitan will be DOE’s third exascale-class supercomputer, following Argonne National Laboratory’s "Aurora" and Oak Ridge National Laboratory’s "Frontier" system. All three DOE exascale supercomputers will be built by Cray utilizing their Shasta architecture, Slingshot interconnect and new software platform.
Link to the congressional report, https://www.energy.gov/sites/default/files/2022-03/FY%202022...Link to the press release, https://www.llnl.gov/news/doennsa-lab-announce-partnership-c...
-
It does seem to me that it's an accurate assessment to say that this project exists because of the nuclear weapons research mandate.
It's not an accurate assessment, in that it's an over-simplification. DoE is a big entity. DoE does weapons research. DoE also does general science research. They are separate responsibilities and have to be, because weapons research is classified and is legally required to be kept separate. The funding pathway that pays for the computing needs of each mission likewise is legally separate and jealously guarded from each other by program managers. The computers are housed, as noted by your quoted press release, in separate laboratories. Those laboratories are run independently of each other, and have distinct responsibilities. Frontier cannot run weapons-related simulations because it's not approved for classified work, which has some really strenuous requirements that the government takes extremely seriously.
It is true to say DoE exists because of nuclear weapons research, that's a big part of their role and history. The distinction is that it's not their only role, and their general science mission is not purely in service of the weapons program legally or practically. It's true to say that the programs coexist together and operate in parallel, but they are separate with their own goals. You're citing the goals of the NNSA computing program, which absolutely are weapons-related, but the Office of Science that funds and operates Frontier, Aurora etc is a separate entity with independent funding and governance within DoE.
The best analogy I can think of is that it's like saying PowerPoint wouldn't exist without Word. Taken very literally there is an element of truth, but it's a gross oversimplification of the actual situation and history.
Note that DoE publicly releases almost all of its unclassified publications on osti.gov for free, including publications using Frontier and the other open supercomputers. OLCF also specifically aggregates publications from researchers using their machines if you wanna get an idea of what is done with them https://www.olcf.ornl.gov/publications/ . You should be able to get the full texts on OSTI.
Would you mind humoring me?
—
For me, this is the fruit of a tree and a part of a long chain of causality. We start with von Neumann fiddling with machines at Los Alamos and wind up at the Accelerated Strategic Computing Initiative from the 90s https://www.ncbi.nlm.nih.gov/books/NBK44974/ and then fast forward to the excascale program in the early-to-mid 2010s.
I am not a professional, but this is an area of interest for me. And I've been tracking it for some time, and in the earlier communications, the ORNL etc seemed to be fairly clear on who was cutting the checks, from 2013,
The Department of Energy’s (DOE) Office of Science and the National Nuclear Security Administration (NNSA) have awarded $25.4 million in research and development contracts to five leading companies in high-performance computing (HPC) to accelerate the development of next-generation supercomputers.
Under DOE’s new DesignForward initiative, AMD, Cray, IBM, Intel Federal and NVIDIA will work to advance extreme-scale, on the path to exascale, computing technology that is vital to national security, scientific research, energy security and the nation's economic competitiveness.
“Exascale computing is key to NNSA’s capability of ensuring the safety and security of our nuclear stockpile without returning to underground testing,” said Robert Meisner, director of the NNSA Office of Advanced Simulation and Computing program. “The resulting simulation capabilities will also serve as valuable tools to address nonproliferation and counterterrorism issues, as well as informing other national security decisions.”
From: https://www.nersc.gov/news-publications/nersc-news/nersc-cen...The program was funded fully in 2016 after many years of partial funding, ORNL's report,
The mission of the Exascale Computing Project (ECP) is the accelerated delivery of a capable exascale computing ecosystem to provide breakthrough solutions addressing our most critical challenges in scientific discovery, energy assurance, economic competitiveness, and national security.
As a multi-lab effort, sponsored by the DOE’s Office of Science and National Nuclear Security Administration, the ECP is chartered with the following tasks:
Developing exascale-ready applications and solutions that address currently intractable problems of strategic importance and national interest.
Creating and deploying an expanded and vertically integrated software stack on DOE HPC pre-exascale and exascale systems.
Delivering US HPC vendor technology advances and deploying ECP products to DOE HPC pre-exascale and exascale systems.
Delivering exascale simulation and data science innovations and solutions to national problems that enhance US economic competitiveness, change our quality of life, and strengthen our national security.
From: https://www.ornl.gov/content/about-ecpI think the other applications are far more important than the weapons research, fwiw. Just as ARPA-net and the internet it spawned ended up being far more valuable than their initial strategic use case. Even if the important stuff is a side project, just like fusion as an eventual source of energy seems to be for NIF, it's great that it's happening at all.
Hope this makes sense!
Thanks!
DoE has always been a "big tent" of separate missions and I totally understand how it gets confusing. Its predecessor the AEC also had a weird mix of practically orthogonal goals. The non-weapons and non-nuclear missions have gotten a lot bigger in the last 30-40 years and are more or less operated independently of each other. NNSA is also functionally an independent agency within DoE, but it's important to note that they do fund pure science work too, with non-weapons related goals.
A HPC cluster might also have men with guns outside which means they don't need quite the same threat model as a cloud provider (say)
This implies needing to pack all of that compute as closely together physically as possible to minimize latency due to speed-of-light limits, which is substantial in massive compute clusters. It takes a long time for light to get from one side of the cluster to the other from the perspective of a computer.
This, in turn, implies needing sophisticated custom cooling since the power density is very high due to packing so much silicon into such a small footprint.
Another issue is that systems this large are at high risk of producing erroneous results due to the quantity of silicon involved creating a large attack surface for bit flips. It is very expensive to run a large supercomputing job only to find out that the output is incorrect and you'll need to start over, so aggressive mitigation, detection, and correction of sporadic data and compute errors is important.
You could run all of these codes on a stack of beige boxes. It would take so long for the computation to complete that it would be extraordinarily expensive in both time and money. These systems focus on codes that would be effectively computationally intractable if you tried to run them in, say, AWS.
such as?
The cloud has many advantages, but high-quality inter-node bandwidth and topology isn’t one of them. In HPC, the network is the most important part of the system.
That said, multi-terabyte memories won’t solve interesting problems; we already had that. When I was working on this 15-ish years ago, the real-world data models had trillions of vertices, never mind edges. And that has only gotten larger with time. A lot of the research ended up focusing on the problem of how do you boil the ocean selectively and incrementally to optimize throughput.
There is no way to trivially throw hardware at the problem; graph-cutting is hard, and you have to do it even within single servers. Even with sophisticated latency-hiding, it ends up being about effective bandwidth in a context where caches are almost useless.
For graph analysis specifically, we could do a lot with big servers, this is true. But it would require a completely different software architecture to the way most graph analysis is done now. This is perpetually on my “copious spare time” lists of projects because there is a big gap here.
Unfortunately, this had two practical problems. First, it turns out that developers are quite poor at reasoning about control flow in latency-hiding architectures generally. It is analogous to reasoning about very complex lock graphs but worse. Second, you still need prodigious quantities of high-quality network bandwidth and topology, even if latency matters much less, for typical HPC problems. At which point you are half the way to a traditional HPC network anyway.
There is still a space for this research in non-HPC applications (like join parallelization in databases), but for HPC the cost-benefit ratio pushes everyone to purpose-built networks. Less learning curve for the devs and you get the exceptional bandwidth and topology the codes need anyway.
https://aws.amazon.com/blogs/aws/now-available-elastic-fabri...
[1] https://www.doeleadershipcomputing.org/wp-content/uploads/20...
The bandwidth and communications are always huge on these supercomputers. That's really the difference between a typical cluster and a supercomputer... The interconnect.