Check these talks from the recent ISC EXACOMM workshop if you want to see why HPC machines and HPC computing are an entirely different league compared to traditional data center computing: https://www.youtube.com/watch?v=9PPGvqvWW8s&list=WL&index=9&... https://www.youtube.com/watch?v=q4LkF33YMJ4&list=WL&index=7
It has Slinghshot-11[1] as interconnection having a raw power of 200GB speed, plus caching and other heavy optimizations.
It is not only the gpu instances but the way it interconnects. This model has even containers available for use.[2]
It is more open.
[1] - https://www.nextplatform.com/2022/01/31/crays-slingshot-inte...
[2] - https://www.lumi-supercomputer.eu/may-we-introduce-lumi/
Why do not DCs appear? Because they have not submitted benchmarks and power measurements.
This is really the key: a supercomputer has the (software) facilities that makes it possible to launch one coordinate job that runs across all nodes. A data centre is just a bunch of computers placed next to each other, with no affordances to coordinate things across them.
At on point in time the hardware differences were much greater between the two, but the fundamental distinction where a supercomputer really is concerned with having the ability to be "one" computer remains.
The rest kind of follows from that, like how a supercomputer that consists of multiple computers needs a fast, low-latency interconnect between them to coordinate and exchange results, while computers in a DC care a lot less about each other.
On the other hand the distinction is fluid. Google could call the indexers that power their search engine a supercomputer, but they prefer to talk about datacenters
What's interesting is that over time, the datacenter folks ended up adding supercomputers to their datacenters, with very large and fast database/blob storage/data warehousing systems connected up to "ML supercomputers" (like supercomputers, but typically only do single precision floating point). The two work well together so long as you scale the bandwidth between them. At the end of the day, any interesting data center has obscenely complex networking technology. For example, TPUs are PCI-attached devices in Google data centers; they plug into server machines just like GPUs. The TPUs themselves have networking between TPUs, that allows them to move important data, like gradients, between TPUs, as needed to do gradient descent and other operations, but the hosts that the TPUs are plugged into have their own networks. The TPUs form a mesh- the latest TPUs form a 3D mesh, but physically implemented through a complex optical switch, while the hosts they are attached to multiple switches which themselves from complex graphs of networking elements. When running ML, part of your job might be using the host CPU to read in training data and transform it, keeping the network busy, keeping some remote disk servers busy, while pushing the transformed data into the TPUs, which then communicate internal data between themselves and other TPUs, over an entirely distinct network. Crazy stuff.