Intel announces the Aurora supercomputer has broken the exascale barrier
intel.com
intel.com
In the November 2023 list, Aurora was also in second place, with an Rmax of 585.34 PetaFLOPS/second.
See https://www.top500.org/system/180183/ for the specs on Aurora, and https://www.top500.org/system/180047/ for the specs on Frontier.
See https://www.top500.org/project/top500_description/ and https://www.top500.org/project/linpack/ for a description of Rmax and the LINPACK benchmark, by which supercomputers are generally ranked. The Top 500 list only includes supercomputers that are able to run the LINPACK benchmark, and where the owner is willing to publish the results.
The jump in Aurora's Rmax scope is explained by Aurora's difficult birth. https://morethanmoore.substack.com/p/5-years-late-only-2 (published when the November 2023 list came out) has a good explanation of what's been going on.
https://ourworldindata.org/grapher/supercomputer-power-flops
That being said, note that the software is also different on the two computers.
AMD's current architecture is very power responsible, and Intel has more or less used watt overfeeding to catch back in performance.
I would say that's a bit more than process efficiency?
IIRC the numbers I've read are that (at least desktop) Intel CPUs should be using something like 0.2 W package power at idle if the OS is correctly configured, regardless of whether it's a performance (K) or "efficiency" (T) model. Most power usage is the rest of the system.
They both have similar frequency and voltage scaling algorithms at this point. You will probably not see 0.2W idle though, both probably idle around 10W on desktop and 5W on laptop. But Intel is getting much more aggressive with "turbo boost" to try to hide their IPC/process deficit vs. AMD/TSMC, to the point that a 14900k will use 120W+ to match the performance of a 7800x3d at 60W.
* Intel is the best for idle (there's several people that have systems that run at less than 5 W for the full system using modified old business minipcs off ebay). Allegedly someone has a 9500T at less than 2 W full system power.
* It doesn't matter which Intel processor you use; all of them for many years will get down to 1 W or less for the CPU at idle. A 14900K will idle just as well as an 8100T, which will be much better than a Ryzen 7950X.
* AMD pretty much never gets below 10 W with any of the Ryzen chiplet CPUs. Only their mobile processors can do it, but they don't sell them retail and they're usually (always?) soldered.
* Every component except the CPU is more important. Your motherboard and PCIe devices need to support power management. You need an efficient PSU (which has nothing to do with the 80-plus rating, which doesn't consider power draw at idle). One bad PCIe device like an SSD or a NIC can draw 10s of watts if it breaks sleep states. Unfortunately, this information seems to be almost entirely undocumented beyond these crowdsourcers.
For a usually idle home-server, Intel seems to be better for power usage, which is unfortunate because AMD tends to have more IO and supports ECC.
[0] https://www.hardwareluxx.de/community/threads/die-sparsamste...
These systems are used for training which is VERY rarely INT8. On Frontier, for example, it's recommended to use bfloat16 or float32 if that doesn't work for you/your application.
Nvidia has FP8 with >=Hopper and supposedly AMD MI300 has it as well although I have no experience with the MI300 so I can't speak to that.
What do those kinds of upgrades entail from a hardware side? Software side? Is this just a horizontal scaling of a cluster?
See the last paragraph of my post for a link to more info.
Is my understanding correct? If yes, then why is it important to build supercomputers with more and more compute? Wouldn't it be better to build smaller systems that focus more on power/cost/space efficiency?
Also, I guess I'm not sure what you mean by "smaller systems that focus more on power/cost/space". A proper queueing system generally efficiently allocates the resources of a large supercomputer to smaller tasks, while also making larger tasks possible in the first place. And I imagine there's somewhat an efficiency of scale in a large installation like this.
There are, of course, many many smaller supercomputers, such as at most medium to large universities. But even those often have 10-50k cores or so.
(In general, efficiency is a consideration when building/running, but not of using. Scientists want the most computational power they can get, power usage be damned :) )
edit: A related topic is capacity vs. capability: https://en.wikipedia.org/wiki/Supercomputer#Capability_versu...
No point in staying up waiting for a job, it'd get rescheduled in the early morning at best.
It wasn't the largest cluster around, IIRC 768 quad-core nodes, but I'm sure the meteorological department would find a way to utilize any extra capacity, so still requiring the whole thing all night.
An IBM/360 has laughably less compute than your phone.
Yes
> why is it important to build supercomputers with more and more compute
A mix of
- research in distributed systems (there are plenty of open questions in Concurrency, Parallelization, Computer Architecture, etc)
- a way to maintain an ecosystem of large vendors (Intel, AMD, Nvidia and plenty of smaller vendors all get a piece of the pie to subsidize R&D)
- some problems are EXTREMELY computationally and financially expensive, so they require large On-Prem compute capabilities (eg. Protein folding, machine learning when I was in undergrad [DGX-100s were subsidized by Aurora], etc)
- some problems are extremely sensitive for national security reasons and it's best to keep all personnel in a single region (eg. Nuclear simulations, turbine simulations, some niche ML work, etc)
In reality you need to do both, and planners know this fact, and have known this fact for decades
But also, bigger systems have more opportunities to achieve higher utilization than smaller systems due to the dynamics of bin packing problem.
The OLCF Frontier user guide[1] has some information on scheduling and Frontier specific quirks (very minor).
Current status of jobs on Frontier:
[kkielhofner@login11.frontier ~]$ squeue -h -t running -r | wc -l
137
[kkielhofner@login11.frontier ~]$ squeue -h -t pending -r | wc -l
1016
The running jobs are relatively low because there are some massive jobs using a significant number of nodes ATM.
[0] - https://slurm.schedmd.com/documentation.html
[1] - https://docs.olcf.ornl.gov/systems/frontier_user_guide.html
EDIT: I give up on HN code formatting
Just FYI: https://news.ycombinator.com/formatdoc
> Text after a blank line that is indented by two or more spaces is reproduced verbatim. (This is intended for code.)
[kkielhofner@login11.frontier ~]$ squeue -h -t running -r | wc -l
137
[kkielhofner@login11.frontier ~]$ squeue -h -t pending -r | wc -l
1016Oh the irony of using Frontier but not "understanding" HF formatting ;).
Supercomputer admins would love to have a single code that used the whole machine, both the compute elements and the network elements, at close to 100%. In fact they spend a significant fraction on network elements to unblock the compute elements, but few codes are really so light on networking that the program scales to the full core count of the machine. So, instead they usually have several codes which can scale up to a significant fraction of the machine and then backfill with smaller jobs to keep the utilization up (because the acquisition cost and the running cost are so high).
Supercomputers have limited utility- beyond country bragging rights, only a few problems really justify spending this kind of resource. I intentionally switched my own research in molecular dynamics away from supercomputers (where I'd run one job on 64-128 processors for a 96X speedup) to closet cllusters, where I'd run 128 indpendent jobs for a 128X speedup, but then have to do a bunch of post-processing to make the results comparable to the long, large runs on the supercomputer (https://research.google/pubs/cloud-based-simulations-on-goog...). I actually was really relieved when my work no longer depended on expensive resources with little support, as my scientific productivity went up and my costs went way down.
I feel that supercomputers are good at one thing: if you need to make your country's flagship submarine about 10% faster/quieter than the competition.
The GPUs used to train them only existed because the DoE explicitly worked with Nvidia on a decade-long roadmap for delivery in it's various supercomputers, and would often work in tandem with private sector players to coordinate purchases and R&D (for example, protein folding and just about every Big Pharma company).
Hell, the only reason AMD EPYC exists is for the same reason.
I know the DOE/Nvidia history quite well as the Chief Scientist of NVIDIA visited LBL around 2005(6? 7?) and talked about their new hardware they were just starting to build and sell, with the goal of getting them into supercomputers.
We asked if they had double precision performance yet (because that was a must for many supercomputer jobs), but at the time, nvidia DP was still lagging SP (I guess it still does?) and we also quibbled about their non-compliance with some esoteric details in IEEE 754. The best part of the whole talk was when he walked us through the idea of visualizing our operations by drawing the matrices as textures, because you can easily see the NaNs- they render as nvidia Green!
I left DOE (Berkeley Lab) shortly after to work in industry because it was clear that ML wasn't going to be innovated in the government labs.
The "DOE made NVIDIA" myth is a story I haven't seen pushed outside the DOE complex. It is true that the supercomputers the DOE pushes could be considered industry subsidies, by providing industry companies with a steady customer with a very high tolerance for unfinished products. That applies to NVIDIA, AMD, Intel, HPE/Cray and IBM more or less equally.
I also want to stress what often gets overlooked: supercomputers are hell to operate and use. Aurora runs on Slingshot, a Cray interconnect. Those things look good on paper. Examples: Cray Aries (and "network quiesces") or Cray DataWarp. Who knows how Slingshot actually works in practice, for a hero run it only needs to hold things together for a few hours. As long as you get a high TOP500 ranking, a supercomputer is a success.
There is no market for those things anymore and they are beholden to the same economics as everything else, hence codes that can't afford an army of PostDocs to work around the bugs and design decisions that are only necessary due to scale of those systems are better suited to plain old mid-range clusters. And I haven't even mentioned the eccentric userland of supercomputers.
There are many reasons the DOE affords to run those behemoths. Some more trivial and petty than most people would like to believe. Like the author of the parent post, I have come to believe that the best bang for the buck on scientific output can be found elsewhere.
The best bang for the buck is never at the very top. The top is just for the biggest bang.
Do you have a source for this claim? Isn't eg. an H100 basically just a RTX GPU with more and faster memory? (Or, at least, an RTX GPU with the same VRAM as an H100 would perform similarly.) And these GPUs were created to run video games. Unless you are referring to something like NVLink?
Yep
> an H100
H100 NVL is the SKU for HPC.
Or doing Numerical Weather Prediction. :-)
But seriously, as a cluster sysadmin, the “128 jobs, followed by post-processing” is great for me, because it lets those separate jobs be scheduled as soon as resources are available.
> expensive resources with little support
Unfortunately, there isn’t as much funding available in places for user training and consultation. Good writing & education is a skill, and folks aren’t always interested in a job that has is term-limited, or whose future is otherwise unclear.
https://hpc.llnl.gov/documentation/tutorials/introduction-pa...
This is a very detailed free book focusing on programming:
But HPC is very diverse. Some care about compute performance, others about memory bandwidth and others about IO performance. Some run a ton of small jobs while others run a single large job.
Since modern warheads are all fusion-type warheads, there's also the fusion stage to consider with even more highly classified top-secret sauce. It appears that the conditions for fusion are triggered by radiation pressure, and that likely makes things even more complicated. Now, you need not just a successful supercritical fission event, but one of the right shape(?), timing, and interaction with other secret-sauce materials that might have their own degradation curves.
So, rather than simulate one design, now you need to simulate hundreds to thousands to explore the full decay-over-time space. Getting the answer wrong means either very expensive premature warhead refurbishments or a nuclear stockpile that wouldn't work properly.
Cracking them open to have a look may not be a good idea. but leaving them alone for another few decades might not be wise either.
Funny fact: A lot of the nuclear weapons that have been destroyed, have removed them from bombs or rockets. But in many case the warheads were moved into storage. Ready to slap them on a rocket if that should become needed.
The US is replacing its nukes: the new nukes have all new parts except the fissile material. Russia is ahead of the US here and has finished replacing its Soviet-era nukes in this way up to the limit of what they are allowed to deploy under the START treaties. (I.e., they might have some Soviet-era nukes, but if so, they need to stay in storage for Russia to stay in compliance with its treaty obligations.)
US and Russia no longer explode nukes to make sure they still work, which is where simulations using supercomputers come in.
I performed molecular dynamics simulations on the Titan supercomputer at ORNL during grad school. At the time, this supercomputer was the fastest in the world.
At least back then around 2012, ORNL really wanted projects that uniquely showcased the power of the machine. Many proposals for compute time were turned down for workloads that were “embarrassingly parallel” because these computations could be split up across multiple traditional compute clusters. However, research that involved MD simulations or lattice QCD required the fast Infiniband interconnects and the large amount of memory that Titan had, so these efforts were more likely to be approved.
The lab did in fact want projects that utilized the whole machine at once to take maximum advantage of its capabilities. It’s just that oftentimes this wasn’t possible, and smaller jobs would be slotted into the “gaps” between the bigger ones.
It took 15-20 years to reach this point [0]
A lot of innovations in the GPU and distributed ML space were subsidized by this research.
Concurrent and Parallel Computing are VERY hard problems.
[0] - http://helper.ipam.ucla.edu/publications/nmetut/nmetut_19423...
The sound barrier was relevant because there was a significant physical effects to overcome specifically when going trans-sonic. It wasn't a question of just adding more powerful engines to existing aircraft. They didn't choose the sound barrier because its a nice number, it was a big deal because all sorts of things behaved outside of their understanding of aerodynamics at that point. People died in the pursuit of understanding the sound barrier.
The 'exascale barrier', afaict, is just another number chosen specifically because it is a round(ish) number. It didn't turn computer scientists into smoking holes in the desert when it went wrong. This is an incremental improvement in an incredible field, but not a world changing watershed moment.
Data, software stack, I/O have suddenly become bottlenecks in multiple places. So yes, it is a watershed moment..
See this for example. Different applications have different scales at which they reach similar problems. Exascale is decent upper bound for most of the fields. If you really dig in deep, the bound may be found at slightly lower value > 500 Pflops. But it's a good rule of thumb to consider 1 EFlop/s to be safe.
Also see this https://irp.fas.org/agency/dod/jason/exascale.pdf
For example, FPGAs were considered as much more viable for sparse matrix computation instead of GPUs, BLAS implementations were not as robust yet, parallel programming APIs like CUDA and Vulkan were in their infancy, etc.
Just because you didn't do well in your systems classes or you think spinning up an EC2 instance on AWS is "easy" doesn't mean it's an easy problem.
That's like saying Einstein or Planck are dummies because AP Physics E&M students can handle basic relatively and quantum theory.
Please leave personal attacks out of this, it is not in the spirit of HN, or in helping to see people's perspectives.
I'm not saying its not a hard problem, not at all. I respect the hell out of the work that has been done here.
I'm saying that the sound barrier is called a barrier for a very good reason. Aerodynamics on one side of the sound barrier are different than aerodynamics on the other. It is a different game entirely. That is why it is considered a barrier. A plane that has superb subsonic aerodynamics will not perform well on the other side of that barrier
Exascale computers on the other hand, while truly amazing, are not operating differently by hitting 10^18 FLOPS. If you're computer does 10^18 -20 FLOPS it is not operating in a fundamentally different set of rules than one running above the exaflop benchmark.
I never said that the achievement wasn't laudable. I argued that there is no barrier there.
If I'm wrong I would appreciate you explaining why doing things at 10^18 FLOPS is fundamentally different than computing just below that benchmark.
Each jump in flops by a magnitude of 10^3 is a significant problem in concurrency, IO, parallelism, storage, and existing compute infrastructure.
Managing racks is difficult, managing concurrent workloads is difficult, managing I/O and storage is difficult, designing compute infra like FPGAs/GPUs/CPUs for this is difficult, etc.
That's kind of my point, an Exaflop is a benchmark and not a 'barrier'. The sound barrier wasn't a 10^3 change in speed, the difference between a subsonic plane and a supersonic plane is measured in percentages when it comes to speed, e.g. a plane that is happy at .8 Mach for a top speed is going to be designed under a completely different set of rules than one that tops out at 1.2 Mach.
Again, I'm not saying that the accomplishments are insignificant, I'm just arguing press-release semantics here.
What you're doing is the equivalent of asking why do we use Gigabyte or Terabyte as a metric.
Just reaching 10^12 floating point operations per second was not something that was done until 2018.
(also please don't be rude or condescending, it detracts from your argument)
It's functionally the same as arguing that GB or TB are arbitrary units to represent storage.
I used to work with/for NERSC and was around when they announced exascale as a target, and now they want to target zettascale. There is no magic threshold where science simulations suddenly work better as you scale up. It's mainly about setting goals that are 10-15 years away to stimulate spending and research.
We most likely crossed paths. How close were you to faculty in AMPLab?
> It's mainly about setting goals that are 10-15 years away to stimulate spending and research.
And that's not a barrier to you?
That feels very condescending about applied research or the amount of effort put into the entire Distributed Systems field.
How close was I to faculty in AMPLab? Pretty close; I attended a retreat one year, helped steer funding their way, tried to hire Matei into Google Research, and have chatted with Patterson extensively around the time he wrote this (https://www.nytimes.com/2011/12/06/science/david-patterson-e...) then later when he worked on TPUs at Google.
(I'm not a stellar researcher or anything, really more of a functionary, but the one thing I do have is an absolutely realistic understanding of how academic and industrial HPC works)
This is not condescending to the field at all. Crossing an arbitrary goal that is very difficult to get to is still impressive. Just stop using the “breaking the barrier” phrase.
> And that's not a barrier to you?
In what way is that a barrier? Barrier and goal aren't synonymous, it's just marketing speak that confuses them.
Running a marathon is not a barrier, it's a goal, even if many people can't reach it. A combat zone at the halfway point of the marathon is a barrier because it requires a completely different approach to solve.
For certain problems there are absolutely magic thresholds. ML was famously abandoned for 2 decades because the computers were too slow, and the ML revolution has only been possible because of having a ~teraflop on a per-machine level. Weather and climate models are another where there have been concrete compute targets, a whole earth model requires about 10 teraflops (hence the earth simulator super computer). ~1 meter resolution is an exascale level target.
We need a benchmark to delineate between large magnitudes.
Furthermore, there are very real engineering problems that had to be solved to even reach this point.
A lot of noobs take GPUs, Scipy, BLAS, CUDA, etc for granted when in reality all of these were subsidized by HPC research.
The DoE's Exascale project was one of the largest buyers for compute for much of the 21st century and helped subsidize Nvidia, Intel, AMD, and other vendors when they were at their worst financially.
Now granted, the original rendition of this saying ("Breaking the sound barrier.") is also arbitrary because mach 1 is the speed of sound travelling through air on planet Earth, it's still a valid question that you did not answer.
At Mach numbers above 1 the compressibility of the air is entirely different. The medium in which the airplane operates behaves differently, in other words. The sound barrier was a barrier because the planes they were using stopped behaving predictably at Mach > 1. They had to learn to design planes differently if they wanted to fly at those speeds.
Mach 1 is an external constraint mandated by the laws of physics. There is a good reason that sound can't travel faster.
That is why it is a barrier to be broken. It is a paradigm shift imposed entirely by the properties of our physical world.
It's an easy to understand mark on a very smooth difficulty curve. Not a barrier.
I rarely do these days.
Shame, really, HPC could use a little bit of high stakes adventure to make it sexier. (Funny how risking death makes things more attractive ?)
There’s something about working with equipment where the line between top performance and a smoking hole is a matter of degree.
Also opens up a lot more Netflix production opportunities and I bet code safety would get a bump as well.
Because of that, I felt a bit of nostalgia when I first saw consumer-accessible GPUs hitting the 1 TFLOP performance level, which now I suppose qualifies as a cheap iGPU.
Unlike the sound barrier in flight, it was for a while thought impossible to fly faster than sound and there are indeed models of subsonic flight that have an infinity in them at the speed of sound. It indeed took new models and radically different designs to fly faster than sound.
Making this computer do a certain number of calculations didn’t involve any such problem to overcome.
While I'm not sure exascale was something like the sound barrier, I do know a lot of hard work has been done to be able to efficiently utilize such large clusters.
Especially the interconnects and network topology can make a huge difference in efficiency[1] and Cray's Slingshot interconnect[2], used in Aurora[3], is an important part of that[4].
[1]: https://www.hpcwire.com/2019/07/15/super-connecting-the-supe...
[2]: https://www.nextplatform.com/2022/01/31/crays-slingshot-inte...
It took significantly longer than it should have if it was just business as usual: "At a supercomputing conference in 2009, Computerworld projected exascale implementation by 2018." [1]
We got the first true exascale system with Frontier in 2022.
Part of the problem was the power consumption and having a purely CPU based system online for an exascale job. From slide 12 from[2]: "Aggressive design is at 70 MW" and "HW failure every 35 minutes".
[1]: https://en.wikipedia.org/wiki/Exascale_computing [2]: https://wgropp.cs.illinois.edu/bib/talks/tdata/2011/aramco-k...
Cool. Please ask it to sue for peace in several parts of the world, in an open way. Whilst it is at it, get it to work out how to be realistically carbon neutral.
I'm all in favour of willy waving when you have something to wave but in the end this beast will not favour humanity as a whole. It will suck up useful resources and spit out some sort of profit somewhere for someone else to enjoy.
A lot of the foundational models used today were trained on Aurora and its predecessors, as well as tangential research such as containerizarion (eg. In the early 2010s, a joint research project between ANL's Computing team, one of the world's largest Pharma companies, and Nvidia became one of the largest customers of Docker and sponsored a lot of it's development)
For example, I have an "AI" project on Frontier. The process was remarkably simple and easy - a couple of Google Meets, a two page screed on what we're doing, a couple of forms, key fob showed up. Entire process took about a month and a good chunk of that was them waiting on me.
Probably half a days work total for 20k node hours (four MI250x per node) on Frontier for free, which is an incredible amount of compute my early resource constrained startup would have never been able to even fathom on some cloud, etc. It was like pulling teeth to even get a single H100 x8 on GCP for what would cost at least $250k for what we're doing. And that's with "knowing people" there...
These press releases talking about AI are intended to encourage these kinds of applications and partnerships. It's remarkable to me how many orgs, startups, etc don't realize these systems exist (or even consider them) and go out and spend money or burn "credits" that could be applied to more suitable things that make more sense.
They're saying "Hey, just so you know these things can do AI too. Come talk to us."
As an added bonus you get to say you're working with a national lab on the #1 TOP500 supercomputer in the world. That has remarkable marketing, PR, and clout value well beyond "yeah we spent X$ on $BIGCLOUD just like everyone else".
That’s why it appears unstoppable.
> Why It Matters: Designed as an AI-centric system from its inception
Announcement was in 2015. I'm curious whether Argonne are pleased with it.
All I'll say is that you can probably guess how they feel about it given the context.
Intel, AMD, and Nvidia were all vendors on Exascale projects [0]
> Intel is dog shit
Intel has issues with execution, but their engineers are still top notch. They are the last American company to actually do semiconductor fabrication, and only fell behind TSMC and Samsung in fabrication only 6-7 years ago because they didn't choose to invest in EUV lithography instead of other methods.
[0] - http://helper.ipam.ucla.edu/publications/nmetut/nmetut_19423...
With all of the taxpayer money they've wasted so far, they could've bought zillions of Cerebras WSE-3 and exceeded 10 exaflops.
"I went to school for > 25 years so my life's work could be selling more ads" isn't a motivator for most of these people.
While both NVIDIA and AMD design their top GPU model for both FP64 and AI/ML workloads, to save on the design cost, you can do AI training using GPUs that have only high AI performance (like RTX 4090 or its workstation counterpart, RTX 6000), without implementing FP64 operations at all (the FP64 performance of RTX 4090 is negligible, being worse than of any decent cheap CPU, it is provided only for compatibility in testing).
Some might, but most of the work done on GPU compute was subsidized by DoE Exascale projects like Aurora, like my anecdote about the DGX product line above.