such as?
such as?
The cloud has many advantages, but high-quality inter-node bandwidth and topology isn’t one of them. In HPC, the network is the most important part of the system.
That said, multi-terabyte memories won’t solve interesting problems; we already had that. When I was working on this 15-ish years ago, the real-world data models had trillions of vertices, never mind edges. And that has only gotten larger with time. A lot of the research ended up focusing on the problem of how do you boil the ocean selectively and incrementally to optimize throughput.
There is no way to trivially throw hardware at the problem; graph-cutting is hard, and you have to do it even within single servers. Even with sophisticated latency-hiding, it ends up being about effective bandwidth in a context where caches are almost useless.
For graph analysis specifically, we could do a lot with big servers, this is true. But it would require a completely different software architecture to the way most graph analysis is done now. This is perpetually on my “copious spare time” lists of projects because there is a big gap here.
Unfortunately, this had two practical problems. First, it turns out that developers are quite poor at reasoning about control flow in latency-hiding architectures generally. It is analogous to reasoning about very complex lock graphs but worse. Second, you still need prodigious quantities of high-quality network bandwidth and topology, even if latency matters much less, for typical HPC problems. At which point you are half the way to a traditional HPC network anyway.
There is still a space for this research in non-HPC applications (like join parallelization in databases), but for HPC the cost-benefit ratio pushes everyone to purpose-built networks. Less learning curve for the devs and you get the exceptional bandwidth and topology the codes need anyway.