HPC is dying, and MPI is killing it (2015)
dursi.ca
dursi.ca
MPI and Spark solve very different problems. The overlap is basically zero and the fact that MPI is flat while Spark fluctuates shows this. HPC is a small small fraction of the job market, the number of folks involved in HPC is tiny compared to all the click-log analyzing Spark and Hadoop programmers.
Both of those systems are bulk parallel systems, with the majority of folks aggregating low information density data. This is not what the MPI clusters are doing modeling weather and calculating subatomic interactions or simulating wind tunnels.
> The idea that the people at Google doing large-scale machine learning problems (which involves huge sparse matrices) are oblivious to scale and numerical performance is just delusional.
Scale is not latency! Google scales to sizes many orders of magnitudes larger than MPI clusters, but it does not run workloads with the same connectivity needs that MPI workloads need. It doesn't run Super Computers, it runs massively parallel bulk embarrassingly parallel computers.
Chapel is a language, MPI is a transport. The author obviously has skills, but they shouldn't be conflated. Chapel can use MPI.
Chapel supports OFI MPI uGNI GASNet. This is not unlike saying don't use SCTP use Python! I am not being charitable.
This isn't exposed to the larger Google, nor to customers. Google itself as an organism doesn't HPC. It is an ad company that runs ad company systems.
Why are there no top-k super computers running on Google?
Please make your point instead of trying to negate whatt I said.
Medium matters a lot- because supercomputers are fundamentally constrained by physics in terms of power dissipation, transistor density, and speed of light, all-optical networks have slightly lower latencies than electrical, and also let you build larger systems (longer cables).
You can read more on paper that my collegues in Google platforms recently published https://arxiv.org/abs/2304.01433 https://cloud.google.com/blog/topics/systems/tpu-v4-enables-...
I am so happy this paper is finally out :-) I led the SRE work during their NPI. It's a pain when you have to wait 3 years to discuss what you worked on.
They are based on OCS, which powers Google datacenter networks https://arxiv.org/abs/2208.10041
Though not in the HPC field myself, I've been in a position of helping researchers working with an HPC cluster. On this hardware something like 16 cores per node was typical, so you run out of scaling room quickly using traditional threading libraries. And you really want to scale beyond this - jobs could take weeks or months otherwise. This means using MPI because that's all the HPC supported. This was a massive pain, because MPI is complicated in ways that are unfamiliar to many researchers, most of whom have no low-level language experience at all.
What the article calls for are HPC techniques that move communication to a lower level API, rather than exposing them to the programmer / end-user. I think that's exactly the right idea for genomics.
Disclaimer: I'm neither an expert in HPC nor genomics.
https://openai.com/research/scaling-kubernetes-to-7500-nodes
So they are basically classic supercomputing workloads.
I do wonder if in hindsight using a Slurm cluster for scheduling, Lustre for data, MPI for any connectivity thats not covered by NCCL, would have been better than trying to make object storage, grpc, kubernetes, ray etc work.
https://openai.com/research/scaling-kubernetes-to-7500-nodes
You're basically saying HPC = MPI users, which is really dismissive of a whole bunch of other people making use of vast compute resources.
I for one can't wait for these masses to convince chip makers that ieee754 floating point really has a lot of crap that we don't need.
The problem is that you do have to support true parallel MPI jobs in those shared clusters though, so MPI just becomes the hammer for everyone else.
Managing the resources a level higher (all resources live in a k8s cluster, slurm under k8s) seems to be the best way to really accommodate both types of loads but most HPC centers are far off from implementing that.
https://slurm.schedmd.com/SC22/Slurm-and-or-vs-Kubernetes.pd...
(I think that presentation has some misconceptions about k8s - most k8s clusters are elastic to a max size - and it sounds like the really want to control most scheduling - but it gives an overview of merging the systems)
Around that time, there was also a period of reservation-based "advanced scheduling" where HPC centers were flirting with making future scheduling promises. The right way to think of these would be like guarantees to get bare metal machine capacity during a certain wall-clock period. In my opinion, the commercial pre-cloud/cloud/virtualization stuff then infected everyone and regressed to time-sharing with fuzzy QoS and lots of over-subscription and dynamic rescheduling.
Of course, these different approaches will all be isomorphic in the end if they explore the full space of application requirements. The traditional paths were just approaching from very different economic priorities. The IaaS folks are incrementally adding more QoS and pricing options which could eventually provide HPC IaaS if carried to full fruition. I.e. future guarantees of significant hardware resources. But as far as I know, those are still in the realm of "talk to a sales rep" and not some automated IaaS request flow at this point.
(I have worked with Slurm, condor, Torque/PBS, gridEngine, DIRAC, and LSF)
Plugging a different scheduler into k8s might be an interesting way of solving this - it seems like there’s a lot of work on scheduler plugins when I last looked. Some of the issues are similar in cloud too - coscheduling by latency.
There’s at least some incentive to not solve this - I remember k8s on Mesos being popular and of course we know how that played out.
Since then … it’s become a bit of a meme, unfortunately. Definitely there still exist workloads assigned to Spark clusters that could run on a laptop, especially if the data happens to be there already. But the space as a whole provides immense value, both enabling jobs that really don’t fit on laptops, and moving the compute for laptop sized jobs to where the data happens to be.
since then, it's become obvious that the 99% of the ML world is using high-level interfaces, blissfully unaware of MPI.
if you want to make a lasting contribution, work on PIM, not MPI ;)