AWS unveils Graviton4 & Trainium2
press.aboutamazon.com
press.aboutamazon.com
This seems ambiguous. Presumably this is 50% more cores per chip. What about "30% better compute performance" and "75% more memory bandwidth": is that per core, or per chip? If the latter, then per-core compute performance would actually be lower.
Also, "up to" could be hiding almost anything. Has anyone seen a source with clearer information as to how per-core application performance compares to earlier Graviton generations?
Of course it might give Amazon room to offer these instances at a lower hourly rate per core, which would ultimately cash out as improved cost/performance for AWS customers.
The performance improvement is on a per-core basis. The pending availability of 96-vCPU Graviton4 instances is icing on the cake!
But you’re right it is ambiguous.
At least in the "old days" there was ( still is ) a secondary market for used server parts..
Don't know how companies like Amazon, Microsoft and Google would frame a question like this so their "green" narratives wouldn't be hurt but I'm sure they'll do an excellent job.
They will run it til the card dies, chuck it in the bin when it does, and until then they can pass some of the savings to users hopefully, win win?
So that's an obvious home for the chips that are no longer available to users.
I had a science project which was cpu bound and it turns out because people bid based on the performance, the old chips end up costing the same in terms of cpu work done/$ (older chips cost less per hr but do less).
aws though was by far the most expensive so switching to like oracle with their ampere arm was a lot cheaper for me.
In 2019, before I left the EC2 Networking / VPC team, we were using M3 instances for our internal services... those machines were probably installed in 2013 or 2014, making them over 5 years old.
With the slowdown in Moore's law and chip speeds, I'd wager that team is still using those M3s now.
Eventually the machines actually start failing, so they need to be retired, but a large portion of machines likely make it to 10 years.
IE. if you replace the 5 year old Xeons with new ultra-efficient ARM chips, wouldn't that save you more power and cooling over x amount of time?
Genuine question.
And you can watch him say it too: https://youtu.be/kHW-ayt_Urk?si=DKyw0-Pk-dhU5zFG&t=323
As a kid I always wanted one of those yellow google search appliances and now you can find them everywhere being used as like lawn ornaments.
Things accessed through network APIs and billed per op or in aggregate. Distributed file systems, databases, even build and regression suite systems.
Another key point is that older generations of servers for full custom cloud environments tend to co-evolve with their environments. The amount of power and cooling for a rack may not support a modern deployment.
Especially if a generation lasts 6 years. You might be able to cascade gen N+1 to N, but N+6 may require a full retrofit. A 6 year old data center that is partially filled as individual servers fail may justify waiting for N+7 or even 8 to cover the cost of the downtime and retrofit.
There is a reason Google announced that they are depreciating servers over 6 years and Meta is at 5 years, vs the old accounting standard of 3 years.
Then of course there is a secondary market for memory and standard PCI cards, but the market for 6 year old tech is mainly spares, so it is unlikely to absorb the full size of the N-6 year data center build.
If you are considering a refurb style resale market for 6 year old tech, it is often the case that the performance per dollar is a non-starter because of the amount of power the older tech consumes.
Hyperscalers design their own datacenter "SKUs" for storage/compute, all the way from power delivery to networking to chassis. These servers are going to be heavily customized and it's unlikely that even if they fit normal form factors that they will work in the same way as COTS devices or things you would buy from Supermicro.
You could possibly make it work. If they sold them. But they don't, and if you're in the market for that stuff, Supermicro will just design it for you anyway, because presumably you have actual money.
And the reality is they're probably either break even or greener doing it this way, as opposed to washing their hands of it and selling servers on Ebay so they can eventually get throw in landfills wholesale by nerds once their startups fail or they get bored of them. Just because you stick your head in the sand doesn't mean it doesn't end up in a landfill.
Is AWS doing that?
Even if Nitro was out of the picture or whatever, and you just had the raw package -- it's not like you can really make a motherboard magically from thin air for these devices based on just the CPU pinout, and the tolerances just for power delivery and memory buses are pretty tight, not to mention a gazillion other things.
More broadly, designing compute that is used purely in-house versus large-scale high-volume COTS designs, through e.g. OEM partners, is literally a difference of years and tens or hundreds of millions of dollars. Support, documentation, supply chain relationships, etc. These take a lot of money to do right, and when you buy servers, part of the purchase goes to those departments, to fund them. Most places are better off just talking to Supermicro if they actually need servers, for that reason. But hyperscalers literally save ridiculous amounts of money by doing it themselves and not doing the other things Supermicro does, like OEM work, support, and NRE on generalist designs that are useful outside to third parties.
Laptops and phones are are already a SoC with IO in a particular form factor, and sever farms will go in the same direction with minor differences in the energy or rack density that come out in the wash.
Yet nobody will ever get noe to play with them in RL. I cant hope to buy one a year from now and stuff it in my home office.
What will all this mean for consumer oriented cpus?
Would it be accurate to say that Intel funds part of the development of consumer cpus with the server cpus? (or is it the other way around). It seems like Xeon chip advances drip downwards after a while.
If AWS and Azure stop buying chips from Intel and AMD, presumably that woud be interesting.
End of Next Year we will get 128 Core Neoverse V4 / Cortex X4 with 3nm. And 3nm Zen 5 EPYC.
Server CPU Market is getting quite exciting.
I'd guess maybe the two directly abutting the core are memory controllers, but maybe they are the stacked memory? Maybe the top and bottom chips are io controllers? It felt like destiny that eventually Nitro was going to be on-package, maybe those are basically big honking nitro-like chips?
For the interconnect I doubt this is their typical interconnect but it doesn't seem completely unreasonable. Even when not running massive clusters they'll still need the interconnect to pair the random collections of machines that people are using.
You don’t watercool 200W chips typically, and you can in theory air cool 8x 800 watt nvidia h100s in a single system. These are also 4-5u systems!
16 chips in one node would be ambitious, I would expect the 16 chip offering to really be several closely located nodes in the same rack/nearby.
I'd expect it to be like Google's TPUs which have 4 "chips" in a "pod", attaching 4 of these pods to a single system doesn't seem unreasonable.
Looking at the corrosponding CPU and RAM of the available instance types it looks like they're using 32-core CPUs in dual socket systems.
(Just like we had to do at microsoft for maia with the sidekick rack of just cooling: https://www.datacenterfrontier.com/machine-learning/article/...)
But I would say that's a pretty small datacenter, wouldn't you?
I guess at the scale those companies are operating it's not that big, but that's still quite a large building !