I remember hearing a lot about rankings of supercomputers, but less so about what they actually achieved.
I remember hearing a lot about rankings of supercomputers, but less so about what they actually achieved.
1. Low latency network, 1-2us. Most servers can't ping their local switch that quickly, let alone the most distant switch for 1M nodes
2. High bandwidth network, at least 200gbit
3. A parallel filesystem
4. Very few node types.
5. Network topology designed for low latency/high bandwidth, things like hypercube, dragonfly, or fat tree.
6. Software stack that is aware of the topology and makes use of it for efficiency and collective operations,
7. Tuned system images to minimize noise, maximize efficiency, and reduce context switches and interupts. Reserving cores for handing interrupts is common at larger core counts.supercomputers do all the hard work in research universities all the time. Hell, astrophysics and research involving telescopes and observatories use em all the time.
Absolutely, they contribute to research all of the time.
Some of them have pages where they list research outputs that they enabled (though this is of course limited to those authors tell them about!).
Besides the outcomes that were adopted by the industry, before cloud computing there was grid computing, exactly to manage such resources at scale.
If you're talking about the more stereotypical "high performance supercomputers", I think that they are still used very liberally within the defense industry. I think Lockheed Martin, for example, uses them for CFD analysis.
1. workload for national labs this is mostly sparse fp64 in my understanding, for warehouse-scale computing is lots of integer work, highly branchy, lots of pointer chasing, stuff like that.
2. latency/reliability vs throughput warehouse-scale computing jobs often run at awful utilization, in the 5-20% range depending on how you measure, in order to respond to shocks of various kinds and provide nice abstractions for developers. fundamentally these systems are used live by humans and human time is very valuable so making sure it stays up always and returns quickly is paramount. In my understanding supercomputing workloads are much more throughput-oriented, where you need to do an enormous amount of computation to get some answer but it doesn't much matter whether the answer comes in one week or two weeks.
3. interconnect warehouse-scale computing workloads are mostly fairly separable and the place where different requests become intertangled is in the database. In the supercomputing world, in my understanding, there are often significant interconnect needs all the time, so extremely high performance networking is emphasized.
Generally given there are a much larger number of less powerful computers, more accessible to much scrappier interests, one would expect more innovation to be done on them.