HPC at Autodesk
forums.autodesk.com
forums.autodesk.com
I'd rather have a real HPC cluster.
And you're still stuck on a non-deterministic high-latency network you can't get rid of, and with very limited hardware configurations.
It's more like a grid than a HPC cluster.
There are only two possible advantages:
- you want a lot of hardware very quickly rather than wait for it to be delivered.
- you don't have the desire/capability to be/hire a network engineer.
(note: I've worked in supercomputing and HPC for over two decades"
The network I was talking about is called UltraCluster which have an extremely high bandwidth and low latency, designed to get great scaling on MPI jobs (as well as ML). Typical instances used with UC are p5, which have 8 H100 nvidia GPUs, 192 vCPUs, 2TB RAM, 3.2Tbps bandwidth PER MACHINE, 900GB/sec between GPU peers, and 8 3.84TB SSDs. They are not marketed as metal instances.
No, it's not like a grid. Your thinking is dated and not representative of how people do HPC on AWS, Azure, or Google.
It seems how people do HPC on AWS is limited by what AWS can do (and maybe costs). Our experience was that even the elastic feature wasn't, and we often couldn't get resources anyway.
Maybe dated, but for context, we had 2TB and 128 real cores a decade ago, and I currently work with Summit-type hardware; I'd rather not admit after how long.
Looking into the Ultraclusters page you linked to in a sibling comment, it seems like the host machines pretty much fill out their PCIe connections with Infiniband networking to reach that figure:
EFA is also coupled with NVIDIA GPUDirect RDMA (P5, P4d) and
NeuronLink (Trn1) to enable low-latency accelerator-to-accelerator
communication between servers with operating system bypass.
https://aws.amazon.com/ec2/ultraclusters/If you care about correct NUMA and HyperThreading usage, and even more so if you care about latency on the CPU (for example for real-time trading), the only things that perform well are either metal of full-machine-but-with-hypervisor.
Even before then, AWS had ways of placing all your jobs on the same rack, which used same-rack switching. Whenever AWS told us stuff about HPC we tested it carefully, and I'd say that before ultracluster, AWS was being misleading about performance, although Azure was not.
Ended up with a bursting type setup where you can use the cloud to spin up more nodes than we have in our cluster, or larger nodes than we have available.
Sounds great. Works okay. The problem was with billing.
Our traditional cluster has been around for a while, upgraded and added on to, and there is a bit of a handshake agreement over who pays how much, based on usage. Every year or two the company departments get together and either ask for less cost or take on more based on usage stats. Each of them pay a fixed cost per year.
With the cloud setup, every job came back with charges from our cloud provider. User C consumed 48 hours of usage with 150 nodes of type Z. That will be $900.
After just a few short months they asked us to turn that capability off.
Were you using savings plans? Were you using spot instances?
In my experience ray in AWS is a good way to badly utilize resources and
waste a lot of money
That pretty much sums up Autodesk's relationship with Amazon in its entirety. OTOH if operations isn't your core competency, it's still money well spent.Edit: Before smashing the vote button I'd talk to someone at Autodesk about just how little oversight goes into their AWS usage and how chaotic the billing is.
The shift towards subscription based bullshit was essentially the start of the effort to oust Bass.
There are also more strategic open source initiatives such as the USD stuff covered here.
Many of us, especially in the research division, would love to put more code out there. (For example some tools that we use internally) The good news on this front is that there is now a sanctioned process for this to happen, and the attitude seems much warmer than when I joined a decade ago.
I’m personally involved in trying to open source some of my own work in the robotics domain, and have been pleasantly surprised with the response.
We'll blog more about this soon but you can certainly give it a try today! https://github.com/outerbounds/metaflow-ray