Amazon Unveils Graviton4: A 96-Core ARM CPU with 536.7 GBps Memory Bandwidth
anandtech.com
anandtech.com
So, uh, 75% more bandwidth but 200% more cores is a net reduction in available bandwidth per core. If you rented 32 cores of a server before, you got a certain amount of bandwidth. If you now rent the same 32 cores, you'll have about half the realized bandwidth?
Only in applications that are constrained by memory bandwidth. Many applications are not constrained by the memory bandwidth on modern architectures with significant cache.
If you have an application constrained by memory bandwidth, carefully selecting the server would make sense.
> In related news, Apple reduced memory bandwidth in their M3 chips by 25% compared to M1/M2:
Not exactly. M3 Max has the same 400GB/sec memory bandwidth as the top end M1 and M2 chips. It’s only certain lower tiers that have less memory bandwidth than their equivalent tiers in previous gens.
But it probably doesn’t matter for most applications.
Sure but I think OPs point is that memory bandwidth is the bottleneck in more applications that you think.
Cache only helps if you're accessing the same data over and over.
My gut instinct is the amount of applications that do a lot of processing but only on a small amount of data are in the minority. The opposite seems a lot more common.
It's also not as easy as GB/s/core, since cores aren't entirely uniform, and data access may be across core complexes.
The work I do could be called data science and data engineering. Outside some fairly trivial (or highly optimized) sequential processing, the CPU just isn't fast enough to saturate memory bandwidth. For anything more complex, the data you want to load is either in cache (and bandwidth doesn't matter) or it isn't (and you probably care more about latency).
After some digging, I've realized that one had 8x8GB ram modules and the slower one had 2x32GB.
I did some benchmarking then and found that it really depends on the workload. The www app was 50% slower. Memcache 400% slower. Blender 5% slower. File compression 20%. Most single-threaded tasks no difference.
The takeaway was that workloads want some bandwidth per core, and shoving more cores into servers doesn't increase performance once you hit memory bandwidth limits.
Anyway, I have near-zero experience in this area so I'm mostly just posting this hoping for someone to explain why I'm wrong.
simdjson (impressive as it may be) existing as an interesting project seems likely to imply that most JSON parsing isn't done with such performant methods on "the typical REST server".
EDIT: Seeing that it's used in many projects including Node.js shifts me back to thinking more highly of the claim that memory bandwidth is becoming the ultimate spec!
I doubt that. The easily way to find out is to disable the L3 cache on your CPU. If your theory is right, the performance drop should be minimal on most applications.
Note: Even with the L3 cache disabled there is still the L1 and L2 caches but those are pretty small.
One of the interesting things for me is how memory "operations" (which is to say transactions of the memory controller) can completely obliterate "bandwidth."
In the past (not sure how true this is on current microarchitectures) the controller between cache and DRAM worked by opening a "page" of dynamic memory. That was an artifact of how DRAMs use both a column address select and a row address select and then multiplex the address bits. So to "read" memory you needed to select the column, then select the row, and then read the memory. The good news was that if you needed the next word of memory in order, you could just read again. And again. Periodically if you were reading the chip would have to ask you to wait while it refreshed its contents.
Anyway, this memory operation of opening a new page was a lot slower than reading the next word in memory. So if you're requests were bouncing all around memory your effective bandwidth was limited by how fast your memory controller could open new pages. That could be one tenth the nominal serial access bandwidth. Controllers had multiple "page" registers so they could hold the state of two (or more) different DIMMS and try to interleave their access across DIMMs to hide the latency aspect of page mechanics.
Generally though, when you get to the point where your measuring memops and trying to layout your physical memory to minimize them you're in a different realm of system optimization.
The emphasis on processor bandwidth on memory even dates back to the CDC Star.
I always find this graph [2] covering decades of hardware evolution fascinating.
[1] https://en.m.wikipedia.org/wiki/Roofline_model
[2] http://www.nextplatform.com/wp-content/uploads/2022/12/donga...
Anyone know some example bytes per flop machine balance numbers for current and eg 20 year old systems? For older ones one source is https://www.cs.virginia.edu/stream/peecee/Balance.html - the "machine balance" for 2003 boxes seems to be between 10 and 17 there. Sadly the "MW/s" unit for machine words is not self-explanatory, what is the word size used.
Intel and AMD both tend to choke bandwidth (memory and PCI lanes) on anything below their server chips.
https://www.anandtech.com/print/17024/apple-m1-max-performan...
https://www.mgt-commerce.com/blog/aws-announces-general-avai...
300 x 1.75 = 525, ~536. Per chip.
If the new chip has more cores, then where is the extra bandwidth hiding?
Edit: Nevermind. A reply confirms the memory uplift is per core.
My guess is that most cloud customers use fairly little compute and low average bandwidth. But everyone wants "dedicated" cores so this makes a ton of sense and Amazon is raking in billions due to stupidity. Remember the recent story how a couple clicks cut the company cloud bill by a shit-ton. That's probably the norm.
The article says R7g is 64 cores for a full CPU and R8g is 96. Where are you getting the 200% from?
That is 192 CPU per instances. I assume AWS will continue to work with AMD and provide 256 vCPU. Would love to see them being compared. ( ARM with Physical CPU core and AMD are SMT thread. )
Also, thank you for sharing the https://www.geektime.co.il/ site, it seems it has original content and you can follow it translated in English.
If you're referring to the general performance of single thread apps between the two yes.
It’s a Neoverse N1 architecture, whereas the new Gravitron is Neoverse N2.
It is an E-ATX form factor, and I can’t tell whether the price makes it a good value for someone who simply wants a powerful desktop rather than ARM-specific testing and validation.
Of course this could all be redesigned to be a desktop PCIe card, but the design assumption that it lives in AWS is literally baked into the silicon.
Never mind the power and cooling requirements. You probably wouldn't appreciate it being next to you while you work.
I would use it to make jerky
This is the tip of the iceberg and all the other zoo of Pytorch primitives also need to be implemented, again on the same hardware, but you get the idea. Never mind the complexity of data movement.
The other Neuron core engine is the piece I work a lot with, the general-purpose SIMD engine. This is a bank of 8x 512-bit-wide SIMD processor cores and there is a general-purpose C++ compiler for it. This engine is proving to be even more flexible than you might imagine.
[1]: https://awsdocs-neuron.readthedocs-hosted.com/en/latest/gene...
Customers will want to migrate anyway because the newest types offer much better performance and pricing. At which point AWS may decide there's no demand for new instances of that type and just remove it from the API.
Selling them? Running other services on them? Just keeping them in reserve? Or all turned into e-waste?
Now, I manage many servers at my job that are much older than that, but only because the operations budget and the capital expenditures budget do not take this into account.
The z16 is designed to run applications that cannot fail. In some cases, the datacenter is expressly built around these systems to accommodate their unique capabilities.
The Graviton line is designed to run as many applications per unit of volume and power possible for a multi-tenant, hyperscale ecosystem.
If a CPU core goes bad on a Graviton chip, you would likely need to redeploy your EC2 instance (or it would be done for you automatically). This would almost certainly have some downtime for your application. In the z16, if an entire CPU goes out, you hypothetically won't drop a single transaction if you followed IBM's guidance.
Mandating that SW/Services companies be forbidden from silicon is every bit as absurd as prevention hardware manufacturers from toching software. The playing field for software is incredibly level these days - everyone uses the same IP vendors, design services & contract fabs.
If you want smaller companies, reform it'll be much easier to get rid of tax pyramiding that favours mega corps.
I'm not arguing that at all. Just that they shouldn't cross so many markets at once.
I think with Apple, Amazon, Microsoft and Google, you have a good competitive landscape to work with right now. All of these players already making their own chips. ~4 solid verticals that are somewhat compatible and highly competitive with each other in large, mostly-overlapping regions.
Honestly, things feel pretty OK to me when you factor it all in. If it was just Microsoft or Apple doing the vertical chip thing and they were going absolutely hockey stick over it, then perhaps. But there are several competent players doing this now.