Is my understanding correct? If yes, then why is it important to build supercomputers with more and more compute? Wouldn't it be better to build smaller systems that focus more on power/cost/space efficiency?
Is my understanding correct? If yes, then why is it important to build supercomputers with more and more compute? Wouldn't it be better to build smaller systems that focus more on power/cost/space efficiency?
Supercomputer admins would love to have a single code that used the whole machine, both the compute elements and the network elements, at close to 100%. In fact they spend a significant fraction on network elements to unblock the compute elements, but few codes are really so light on networking that the program scales to the full core count of the machine. So, instead they usually have several codes which can scale up to a significant fraction of the machine and then backfill with smaller jobs to keep the utilization up (because the acquisition cost and the running cost are so high).
Supercomputers have limited utility- beyond country bragging rights, only a few problems really justify spending this kind of resource. I intentionally switched my own research in molecular dynamics away from supercomputers (where I'd run one job on 64-128 processors for a 96X speedup) to closet cllusters, where I'd run 128 indpendent jobs for a 128X speedup, but then have to do a bunch of post-processing to make the results comparable to the long, large runs on the supercomputer (https://research.google/pubs/cloud-based-simulations-on-goog...). I actually was really relieved when my work no longer depended on expensive resources with little support, as my scientific productivity went up and my costs went way down.
I feel that supercomputers are good at one thing: if you need to make your country's flagship submarine about 10% faster/quieter than the competition.
The GPUs used to train them only existed because the DoE explicitly worked with Nvidia on a decade-long roadmap for delivery in it's various supercomputers, and would often work in tandem with private sector players to coordinate purchases and R&D (for example, protein folding and just about every Big Pharma company).
Hell, the only reason AMD EPYC exists is for the same reason.
I know the DOE/Nvidia history quite well as the Chief Scientist of NVIDIA visited LBL around 2005(6? 7?) and talked about their new hardware they were just starting to build and sell, with the goal of getting them into supercomputers.
We asked if they had double precision performance yet (because that was a must for many supercomputer jobs), but at the time, nvidia DP was still lagging SP (I guess it still does?) and we also quibbled about their non-compliance with some esoteric details in IEEE 754. The best part of the whole talk was when he walked us through the idea of visualizing our operations by drawing the matrices as textures, because you can easily see the NaNs- they render as nvidia Green!
I left DOE (Berkeley Lab) shortly after to work in industry because it was clear that ML wasn't going to be innovated in the government labs.
The "DOE made NVIDIA" myth is a story I haven't seen pushed outside the DOE complex. It is true that the supercomputers the DOE pushes could be considered industry subsidies, by providing industry companies with a steady customer with a very high tolerance for unfinished products. That applies to NVIDIA, AMD, Intel, HPE/Cray and IBM more or less equally.
I also want to stress what often gets overlooked: supercomputers are hell to operate and use. Aurora runs on Slingshot, a Cray interconnect. Those things look good on paper. Examples: Cray Aries (and "network quiesces") or Cray DataWarp. Who knows how Slingshot actually works in practice, for a hero run it only needs to hold things together for a few hours. As long as you get a high TOP500 ranking, a supercomputer is a success.
There is no market for those things anymore and they are beholden to the same economics as everything else, hence codes that can't afford an army of PostDocs to work around the bugs and design decisions that are only necessary due to scale of those systems are better suited to plain old mid-range clusters. And I haven't even mentioned the eccentric userland of supercomputers.
There are many reasons the DOE affords to run those behemoths. Some more trivial and petty than most people would like to believe. Like the author of the parent post, I have come to believe that the best bang for the buck on scientific output can be found elsewhere.
The best bang for the buck is never at the very top. The top is just for the biggest bang.
Do you have a source for this claim? Isn't eg. an H100 basically just a RTX GPU with more and faster memory? (Or, at least, an RTX GPU with the same VRAM as an H100 would perform similarly.) And these GPUs were created to run video games. Unless you are referring to something like NVLink?
Yep
> an H100
H100 NVL is the SKU for HPC.
Or doing Numerical Weather Prediction. :-)
But seriously, as a cluster sysadmin, the “128 jobs, followed by post-processing” is great for me, because it lets those separate jobs be scheduled as soon as resources are available.
> expensive resources with little support
Unfortunately, there isn’t as much funding available in places for user training and consultation. Good writing & education is a skill, and folks aren’t always interested in a job that has is term-limited, or whose future is otherwise unclear.
https://hpc.llnl.gov/documentation/tutorials/introduction-pa...
This is a very detailed free book focusing on programming:
But HPC is very diverse. Some care about compute performance, others about memory bandwidth and others about IO performance. Some run a ton of small jobs while others run a single large job.
No point in staying up waiting for a job, it'd get rescheduled in the early morning at best.
It wasn't the largest cluster around, IIRC 768 quad-core nodes, but I'm sure the meteorological department would find a way to utilize any extra capacity, so still requiring the whole thing all night.
An IBM/360 has laughably less compute than your phone.
The OLCF Frontier user guide[1] has some information on scheduling and Frontier specific quirks (very minor).
Current status of jobs on Frontier:
[kkielhofner@login11.frontier ~]$ squeue -h -t running -r | wc -l
137
[kkielhofner@login11.frontier ~]$ squeue -h -t pending -r | wc -l
1016
The running jobs are relatively low because there are some massive jobs using a significant number of nodes ATM.
[0] - https://slurm.schedmd.com/documentation.html
[1] - https://docs.olcf.ornl.gov/systems/frontier_user_guide.html
EDIT: I give up on HN code formatting
Just FYI: https://news.ycombinator.com/formatdoc
> Text after a blank line that is indented by two or more spaces is reproduced verbatim. (This is intended for code.)
[kkielhofner@login11.frontier ~]$ squeue -h -t running -r | wc -l
137
[kkielhofner@login11.frontier ~]$ squeue -h -t pending -r | wc -l
1016Oh the irony of using Frontier but not "understanding" HF formatting ;).
I performed molecular dynamics simulations on the Titan supercomputer at ORNL during grad school. At the time, this supercomputer was the fastest in the world.
At least back then around 2012, ORNL really wanted projects that uniquely showcased the power of the machine. Many proposals for compute time were turned down for workloads that were “embarrassingly parallel” because these computations could be split up across multiple traditional compute clusters. However, research that involved MD simulations or lattice QCD required the fast Infiniband interconnects and the large amount of memory that Titan had, so these efforts were more likely to be approved.
The lab did in fact want projects that utilized the whole machine at once to take maximum advantage of its capabilities. It’s just that oftentimes this wasn’t possible, and smaller jobs would be slotted into the “gaps” between the bigger ones.
Yes
> why is it important to build supercomputers with more and more compute
A mix of
- research in distributed systems (there are plenty of open questions in Concurrency, Parallelization, Computer Architecture, etc)
- a way to maintain an ecosystem of large vendors (Intel, AMD, Nvidia and plenty of smaller vendors all get a piece of the pie to subsidize R&D)
- some problems are EXTREMELY computationally and financially expensive, so they require large On-Prem compute capabilities (eg. Protein folding, machine learning when I was in undergrad [DGX-100s were subsidized by Aurora], etc)
- some problems are extremely sensitive for national security reasons and it's best to keep all personnel in a single region (eg. Nuclear simulations, turbine simulations, some niche ML work, etc)
In reality you need to do both, and planners know this fact, and have known this fact for decades
Also, I guess I'm not sure what you mean by "smaller systems that focus more on power/cost/space". A proper queueing system generally efficiently allocates the resources of a large supercomputer to smaller tasks, while also making larger tasks possible in the first place. And I imagine there's somewhat an efficiency of scale in a large installation like this.
There are, of course, many many smaller supercomputers, such as at most medium to large universities. But even those often have 10-50k cores or so.
(In general, efficiency is a consideration when building/running, but not of using. Scientists want the most computational power they can get, power usage be damned :) )
edit: A related topic is capacity vs. capability: https://en.wikipedia.org/wiki/Supercomputer#Capability_versu...
Since modern warheads are all fusion-type warheads, there's also the fusion stage to consider with even more highly classified top-secret sauce. It appears that the conditions for fusion are triggered by radiation pressure, and that likely makes things even more complicated. Now, you need not just a successful supercritical fission event, but one of the right shape(?), timing, and interaction with other secret-sauce materials that might have their own degradation curves.
So, rather than simulate one design, now you need to simulate hundreds to thousands to explore the full decay-over-time space. Getting the answer wrong means either very expensive premature warhead refurbishments or a nuclear stockpile that wouldn't work properly.
Cracking them open to have a look may not be a good idea. but leaving them alone for another few decades might not be wise either.
Funny fact: A lot of the nuclear weapons that have been destroyed, have removed them from bombs or rockets. But in many case the warheads were moved into storage. Ready to slap them on a rocket if that should become needed.
The US is replacing its nukes: the new nukes have all new parts except the fissile material. Russia is ahead of the US here and has finished replacing its Soviet-era nukes in this way up to the limit of what they are allowed to deploy under the START treaties. (I.e., they might have some Soviet-era nukes, but if so, they need to stay in storage for Russia to stay in compliance with its treaty obligations.)
US and Russia no longer explode nukes to make sure they still work, which is where simulations using supercomputers come in.
But also, bigger systems have more opportunities to achieve higher utilization than smaller systems due to the dynamics of bin packing problem.