Graviton2 and Graviton3
muratbuffalo.blogspot.com
muratbuffalo.blogspot.com
Turns out memory access speed is more or less the entire game for everything except scientific computing or insanely optimized code. In the real world, CPU frequency seems to matter much less than DRAM timings, for example, in everything but extremely well engineered games. It'll be interesting to learn (if we ever do) how much of the "real-world" 25% performance gain is solely due to DDR5.
I remember getting my AMD K8 Opteron around 2003 or 2004 with the first on-die memory controller. Absolutely demolished Intel chips at the time in non-synthetic benchmarks.
for insanely unoptimized code, such as accidentally ending up writing something compute intensive in pure python, its very plausible for it to be compute constrained -- but less because of the hardware and more because 99 %or 99.9% of the operations you're asking the cpu to perform are effectively waste.
It sounds interesting. https://www.bunniestudios.com/blog/?p=3554
Having active cooling, making the system noisier and potentially hotter, seems like a pretty big down side.
https://www.anandtech.com/show/17047/the-intel-12th-gen-core...
It's memory latency.
The latter might take an order of magnitude more space, while still being faster.
An example of such a problem is the Cuckoo Cycle Proof-of-Work [1].
(1) It encodes and decodes a protocol from a potentially untrusted source, so there is obviously a lot of waiting for previous results. That much is clear, however I expected profilers to show me some causal link between serial nature and slow execution, but they don't. I have tried perf, Valgrind-Callgrind and AMD μProf (because I have a Ryzen CPU on both of my main private computers). I'm not sure if the tools suck, my test cases suck, or I just don't know how to interpret the tools' results - assignment of cost to lines of code seeming highly unreliable is my main problem. Maybe the stupid things (most of optimization is about not doing stupid things, after that it gets properly hard) that these profilers are designed to catch aren't the stupid or unavoidable things my code is doing.
You can find a nice spreadsheet from Intel here:
https://download.01.org/perfmon/TMA_Metrics.xlsx
Also using Intel tools gives you much clearer answers. There are a bunch of scripts on top of perf that try to automate this. You can find them in pmu-tools git repo.
The basic workflow is always the same. Find out which part of the pipeline gets bottlenecked. Is it stuck in PCIE IO, memory, instruction decoding, etc. in my experience most of the time it is memory due to heavy pointer dereferencing and it is just a sad state of programming languages these days.
There are a bunch of memory benchmark tests that demonstrate these effects very well in the NUMA tools package source repository.
The problem with memory being the slow part is that it affects instruction fetch/decode cycle as well.
Poor instruction selection and sequencing can impact throughout because of port saturation (e.g. in case of AVX2) and leave other dispatch ports idle, premature store buffer flushes where you send 1-entry updates instead of sending stuff in chunks, leaving 70% of store buffer bandwidth idle, etc.
There are a lot of foot guns in a modern processor and usually it’s a mixture of these problems with one being dominant. A lot of times it is not possible to completely address the problem without rebuilding the software from scratch and properly leveraging hardware knowledge from the beginning when architecting your code.
Yep, this is exactly the case - also, on systems that are busy and context-switching often and thus flushing their cpu caches more frequently. Combine the two, busy systems running loads of un-optimized code, and boom, you have described how most computers run in the real world. This is why "synthetic" benchmarks, which are well designed code running on quiet machines more or less match up to CPU Frequency exclusively.
I don't really have any good charts to show you, but you might checkout an old review of the processor I mentioned as having one of the first on-die memory controllers: https://techreport.com/review/5655/amds-opteron-146-processo...
The AMD Opteron 240 1.4GHz keeps up with chips close to 2x it's frequency - and the memory access times are close to 1/2 as costly (ie: almost all the performance gain from 2x frequency is made up by 1/2 memory access time) - this makes sense, but remember these are well optimized applications (POV-Ray and Lightwave were extremely synthetic). In the real world, opening 10 misc windows applications from 2003, the K8 (particularly when overclocked) was a _beast_.
Honestly in cases you mention, badly designed processes killing the Cpu, I fail to see how faster ram makes a huge difference.
So basically any program written in a language with pointer types exclusively.
In theory the GC can also rearrange memory to 'compact' it. I'm not aware of this optimization in practice.
First, most real unoptimised code faces many issues before memory bandwidth. During my PhD, the optimisation guys doing spiral.net sat nextdoor and they produced beautiful plots of what limits performance for a bunch of tasks and how each optimisation they do removes an upper bound line until last they get to some bandwidth limitation. Real code will likely have false IPC dependencies, memory latency problems due to pointer chaising or branch mispredictions well before memory bandwidth.
Then the database workload is something I would consider insanely optimized. Most engines are in fierce performance competition. And normally they hit the memory bandwidth in the end. This probably answers why the author is not comparing to EPYC instances that have the memory bandwidth to compete with Graviton.
Then the claims that they choose not to implement SMT or to use DDR5 are both coming from their upstream providers.
And if they designed the CPU wouldn't they decide which memory controller is appropriate? Seems like AWS should get as much credit for their CPUs as Apple gets for theirs.
Bottom line for Graviton is that a lot of AWS customers rely on open source software that already works well on ARM. And the AWS customer themselves often write their code in a language that will work just as well on ARM. So AWS can offer it's customers tremendous value with minimal transition pain. But sure, if you have a CPU-bound workload, it'll do better on EPYC or Xeon than Graviton.
ARM is just a design. AWS brought it to market. ARM-based server processors are still rare on the ground. IIRC Equinix Metal and Oracle Cloud offer them (Ampere chips) but not GCP or Azure.
We've tested Graviton2 for data warehouse workloads and the price/performance was about 25% cheaper and 25% faster than comparable Intel-based VMs. Still crunching the numbers but that's the approximate shape of the results.
[Imagine "you made this, I made this" meme here]
It occurs to me that AWS might have far more insight into "real workloads" than any CPU designer out there. Do they track things like L1 cache misses across all of EC2?
https://www.brendangregg.com/blog/2021-07-05/computing-perfo...
See slide 26 (and the rest ofc :)).
The development world today looks very different. Back then, language support for other architectures was more bespoke and CPU vendors had to add support for their chips. Today, there are plenty of very rich, platform-agnostic (both CPU and OS) libraries. Additionally, mobile development has sufficiently matured ARM development that I don’t think that argument holds. If it did, then developers wouldn’t be able to develop on their x86 MacBooks and deploy to their mobile Apple devices (yes it’s ARM now but it hasn’t been for the majority). I think the plain x86-box -> server story is pretty solid for but the cloud has changed that. Everyone is now starting out in the cloud with CPU-agnostic languages where switching architectures usually is as simple as changing 1 line in a config. In some cases it matters but the vast majority of SW dev shops don’t feel this like you used to in the 90s and 00s. Plus M1s now provide developers with local ARM development.
If anything programmers are adopting ARM based computers faster than the rest of the market. As pretty much every developer tool gets ported for Apple silicon every company is going to shrug and go "May as well release an ARM Windows/ARM Linux build as well".
Linus' reasoning makes sense, but the real world disagrees with him (at least in our case).
2. With languages like java, go, python, node it doesn’t even matter
3. Devs are migrating to arm en masse (M1s)