TOP500 at ISC’26: We have a New Number 1 Supercomputer
chipsandcheese.com
chipsandcheese.com
my knowledge is 10+ years out of date, but once upon a time if they'd chosen to, Google could have had _several_ entries in the top 10 of the TOP500 list
It's just poker, they didn't want to tip their hand
(These are the systems to which GP was referring at Google.)
I know Google wants to compare their stuff to El Capitan or whatever but the comparison does not seem valid to me.
There's likely also some need for fusion plasma containment and other related simulations.
Most of the time, it just that it’s a hassle. It takes a while to prep and tune a big hero run for benchmarking, and if you spend a billion dollars on a cluster, it’s making you a lot more than that. Taking it down for a day or two stops the money printers.
Sometimes you want to show off what you can do to dissuade others from fucking with you. Sometimes you want to undersell your capabilities to hide your true ability. Sometimes you want others to think you are underselling your capabilities when you are actually at a disadvantage.
Also you should read the second sentence of the CTBT Wikipedia article to find out why it's not even in force (spoiler: US hasn't ratified it).
Plus one think I like to say is that if a bullet is flying towards you, you could know everything about the chemistry of the gunpowder and the composition of the alloy without it affecting what happens next.
Most companies with huge systems don't participate.
Ironically, the related Graph500 benchmarks reflect this better. Performance is dependent more on using the hardware better than better hardware per se.
The new Chinese supercomputer beats all US supercomputers also in HPCG, not only in Linpack.
What is remarkable is that this was done despite the US attempts of sabotaging HPC in China by "sanctions".
This uses custom CPUs designed in China, which implement an Armv9-A ISA with SME (scalable matrix extension) and which use fast HBM memory. These CPUs are fast enough that they do not need any GPUs for exceeding the throughput of the American supercomputers, which use GPUs. This is like in the Japanese Fugaku, which was the first to implement the Armv8-A ISA with SVE, but which now is rather old.
Like in all CPU-based supercomputers, for this new Chinese supercomputer it is much easier to reach a higher percentage of the theoretical maximum throughput, when solving any problem. So for most practical problems it will be faster than a GPU-based supercomputer that would have the same theoretical maximum throughput.
So this is a much more interesting supercomputer than those built by just buying some HPC racks from HPE (Cray). Because China was forbidden to buy the American equipment, they had to innovate and design their own. Eventually they made something better than what they could not buy.
I don’t have a dog in this fight and I no longer work in HPC. Most modern workloads are severely bandwidth bound. The only aspect of the hardware that matters is bandwidth and that is not materially differentiated. The frontier is scheduler design, which is pure software and difficult computer science. HPC competitions avoid problems with a software solution because it isn’t in their interest as hardware manufacturers.
This result is impressive, sort of, but not in the way people are imagining. I was equally dismissive of the previous leader for the same reasons. For most applications, these benchmarks are legacy pagentry.
I agree with what you say about benchmarks, but that is precisely why the advantage of this supercomputer over the following American supercomputers will be even greater in more demanding workloads than Linpack and HPCG.
It has already shown this by having an advantage in HPCG of greater than 26% over the fastest US system, while in Linpack its advantage is of only 22%.
Thus its position in the top cannot be dismissed as insignificant, because it more likely underestimates than overestimates this system.
Also the programming effort for writing an efficient program will be lower than for the GPU-based US supercomputers.
The memory bandwidth is something you can buy. Exotics were >1 TB/s over a decade ago, so 4 TB/s in 2026 is not that impressive. For all practical purposes, these CPUs are also still exotics, you can’t just buy them. I would be very surprised if the memory bandwidth of US exotics haven’t improved over the last 10-15 years.
In any case, for real workloads scalability is mostly a software theory problem at this point and that is still a dark art without much literature.
The only US-designed CPU "exotics" are the Intel Xeon Max CPU series, which use HBM like the Chinese CPUs, but which have a theoretical maximum throughput of only 1.6 TB/s per socket, i.e. 5 times slower than the new Chinese CPUs.
Moreover, the users of Intel Xeon Max complained that they cannot reach the theoretical memory bandwidth. I do not know if that was due to some bug that might have been solved later by Intel with a microcode update or a new mask set stepping.
The server CPUs with standard DIMMs, which will be launched by AMD and Intel next year, will have a memory bandwidth of around 1 TB/s per socket.
The AMD MI300 GPU used in the fastest US supercomputer has a memory throughput of 5.2 TB/s per socket, so lower than the 8 TB/s per socket of the Chinese CPU, which explains why the advantage of the Chinese system increases in the benchmarks more dependent on memory performance.
The latest AMD Instinct GPU, MI355X, increases the memory bandwidth to 8 TB/s, so equal to the Chinese CPU.
However, it may pass some time until someone will build such a big system with MI355X, though perhaps the existence of this new contender might prompt the US labs to upgrade their systems by replacing the older AMD GPUs with newer AMD GPUs.
This is absolutely going to bite us in the face in five to ten years.
Didn't the DoD at one point build a 1k+ PS3 cluster based on their multi-core chip and had a mini supercomputer CotS?
I remember Sony not liking that people were buying them for other things rather than gaming (iirc they were losing money on hardware at the time) so they bricked linux support soon after.
The Air Force did (Condor) and it hit #33 on the 2010 Top500.
The last time when China had the fastest supercomputer, it was more than 20 times slower than this one and more than 8 times less efficient in energy consumption.
Moreover, its capability was overestimated by the Linpack benchmarks and in other workloads its performance was much less impressive.
For this system, it is the opposite situation. Its result in Top500 underestimates it capability. In other more demanding workloads, where the influence of the memory bandwidth and latency is stronger, its advantage over the US supercomputers is greater than in Top500.
One could theoretically drive home with a ready-to-go rack of non-American and/or non-x86 supercomputer nodes at any point in time across the last few decades, sometimes even with non-NVIDIA/AMD massively parallel coprocessor cards. Nobody did.
If China(or any country) would _ship_ these alternative supercomputer hardware, only then anything could change.
Based on the ARMv9.2.
[1] https://www.nextplatform.com/hpc/2026/06/25/a-deep-dive-on-c...
In the CPU cores designed by the Arm company, SME has been added only in the latest generation of Armv9.3-A CPUs, which was launched last year.
For each level of Armv9, there are many mandatory features and many optional features.
If the Chinese CPU does not implement all the mandatory Armv9.3-A features (and we do not know anything about this), then it will still be considered only an Armv9.2-A CPU, but even in that case it should be referred as an Armv9.2-A + SME, in order to not confuse it with the Armv9.2-A CPUs that have been used for a few years in smartphones, laptops and mini-PCs and which do not have SME, so they cannot have a comparable performance.
I’m sure there is a good reason for this, which is..?
Many systems have the node count to be able to run such benchmarks, but are not optimized or even capable of running them. Having these systems run these large calculations in a sustained way, and performing, is what separates a bunch of nodes together in a data center from an actual cluster that is able to run a benchmark like HPL or HPCG.
To sustain 70 - 80% of peak performance across hundreds of thousands of cores, you need a real low-diameter, high-bandwidth, low-latency fabric and a balanced memory subsystem, running on a system w almost no failures or network issues. A loosely-coupled cluster with an oversubscribed fat-tree will 'run Linpack' and then post an Rmax that's a small fraction of its naive peak.
Also, have a look at the Green500, and the systems there. This is not about bragging rights vendors, this is about placing commodity hardware, tuning it, and bringing it up to health in a way that squeezes all of that last performance possible out of the clusters on those lists. That's the opposite of vendor flexing - it's a craft that you cannot see in a simple node count, as some have been comparing here.
If you ever worked on this field, and with the vendors at this scale, you would know. Its not easy, its actually very hard.
... and imagine you need to deploy this, systems at this scale, w technologies that are sometimes just emerging and sometimes even need proper field testing * every 6 months * to be able to reach the scale and stability to land on these lists.
The reason is that with GPUs it is far more difficult to reach a great percentage of the maximum theoretical throughput. Most GPU programs reach only a very small fraction of what is theoretically possible, and in the best cases one may reach something like 50% to 60% of the maximum.
This CPU-based supercomputer has demonstrated reaching 80% of the theoretical maximum throughput, and this is typical for CPU-based supercomputers. It is much easier to write efficient programs for CPUs.
The new custom Chinese CPUs, which use SME, the Arm Scalable Matrix Extension, are fast enough that they have beaten all GPU-based supercomputers, so there was no need to use GPUs.
Moreover these CPUs use HBM for a very fast memory interface, so in the benchmarks that depend more on memory bandwidth they have an even greater advance over the US GPU-based supercomputers. Thus there really was no point in using GPUs.
GPUs are necessary only when your CPUs are not good enough, which was not the case here.
In the recent past, the Japanese Fugaku used the same approach, of avoiding GPUs. At that time, their custom CPUs using the Armv8-A ISA with SVE were the first which used this ISA in HPC, but now that ISA variant is obsolete in comparison with the Armv9-A ISA with SME, which is implemented in these new custom Chinese CPUs.
If all you need to do is matmuls then you can definitely go past this
Linpack consists mostly of matmuls, but nonetheless there are additional operations that prevent GPUs to reach the high utilization of over 80% that is normal for CPUs, so that a throughput over 50% is considered good at the scale of supercomputers.
At the scale of a supercomputer, the utilization factor is considerably less than for an individual GPU or CPU, because the big matrix is split in blocks and the matrix multiplications are computed on different boards and in different racks, then the results are assembled, so there is a communication overhead.
The former Intel Xeon Phi, with a large number of cores that were weak except for their vector execution units, resembled GPUs in failing to reach a high utilization on Linpack.
Most claims about the cost of emulating FP64 on GPUs are wrong, because they assume that only the significand of floating-point numbers must be extended.
In reality it is even more important to extend the exponent, because with the exponent of FP32 overflows would be much too frequent in scientific/technical computations to accomplish anything.
The minimum FP64 emulation on FP32-capable GPUs requires 3 numbers per emulated FP64, which may be 3 FP32 numbers, or the exponent may be an Int32, if that works better on the target GPU. An emulated FP64 operation is likely to be at least 20 times slower than a FP32 operation.
That is much faster than the 1:64 ratio provided in hardware by an NVIDIA GPU, but even on the fastest FP32 GPUs it is too slow to compete with CPUs, in a professional setting.
FP64 emulation on a GPU can be useful only in a home computer, which may have a rather weak CPU and increasing the FP64 throughput using the GPU can be done at no additional cost, so it can be worthwhile.
DoE compute budgets are ~10B USD across labs. AI training is a trillion-dollar workload. Different league.
The AI oriented GPUs or TPUs have either weak FP64 throughput or they may not support FP64 at all.
They can compete neither with CPUs nor with GPUs that have good FP64 support, like the AMD CDNA datacenter GPUs, which occupy all the top places among American supercomputers.
NVIDIA has stopped improving the FP64 throughput even in their "datacenter" GPUs, abandoning this nowadays smaller market to AMD.
The AMD CDNA GPUs can be used for both HPC and AI, so only an AI cluster based on them could have dual use, but most who want AI choose NVIDIA.
The fastest US supercomputer, El Capitan at Lawrence Livermore National Laboratory, reaches only 1.809 exaflops, while this reaches 2.198, 22% higher.
On the HPCG benchmark, which is more strongly influenced by memory bandwidth, the advance over the GPU-based El Capitan is even greater, of 26.4%.
The OS powering 0% of the 500 supercomputers of the Top 500. But this time, it has to be Windows, right? Amirite?
Ah, no, just kidding: it's "Kylin OS". It used to be a BSD derivative and now it's just based on Linux.
I know, I know: "It's a heavily modified Linux". Whatever, it's not Windows and that makes me very happy.