Graviton 3, Apple M2 and Qualcomm 8cx 3rd gen: a URL parsing benchmark
lemire.me
lemire.me
Also, I'm not sure if "correcting" the numbers for 3 GHz is reasonable and reflects real-life performance. Perhaps some throttling could be applied to test the CPUs using a common frequency?
https://buildjet.com/for-github-actions/blog/a-performance-r...
I wish this were only a theoretical concern, a theoretical incentive, but its not. Github Actions is slow, and Gitlab suffers from a similar problem; their hosted SaaS runners are on GCP n1-standard-1 machines. The oldest machine type in GCP's fleet, the n1-standard-1 is powered by a variety of dusty, old CPUs Google Cloud has no other use for, from Sandy Bridge to Skylake. That's a 12 year old CPU.
But the cloud provider prefers the latter because it has 500% more cores for 50% more power. Which is why the latter still goes for >$2000 and the former is <$15.
It really does not depend on the workload, when those workloads we're talking about are by-and-large bounded to 1vCPU or less (CI jobs, serverless functions, etc). Ice Lake cores are substantially faster than Ivy Bridge; the 8352V will be faster in practically any workload we're talking about.
However, I do agree with this take, if we're talking about, say, lambda functions. The reason being that the vast majority of workloads built on lambda functions are bounded by IO, not compute; so newer core designs won't result in a meaningful improvement in function execution. Put another way: Is a function executing in 75ms instead of 80ms worth paying 30% more? (I made these numbers up, but its the illustration that matters).
CI is a different story. CI runs are only bound by IO for the smallest of projects; downloading that 800mb node:18 base docker image takes some time, but it can very easily and quickly be dwarfed by all the things that happen afterward. This is not an uncontroversial opinion; "the CI is slow" is such a meme of a problem at engineering companies nowadays that you'd think more people would have the sense to look at the common denominator (the CI hosts suck) and not blame themselves (though, often there's blame to go around). We've got a project that can build locally, M2 Pro, docker pull and push included, in something like 40 seconds; the CI takes 4 minutes. Its the crusty CPUs; its slow networking; its the "step 1 is finished, wait 10 seconds for the orchestrator to realize it and start step 2".
And I think we, the community, need to be more vocal about this when speaking on platforms that charge by the minute. They are clearly incentivized to leave it shitty. It should even surface in discussions about, for example, the markup of lambda versus EC2. A 4096mb lambda function would cost $172/mo if ran 24/7, back-to-back. A comparable c6i-large: $62/mo; a third the price. That's bad enough on the surface, and we need to be cognizant that its even worse than it initially appears because Amazon runs Lambda on whatever they have collecting dust in the closet, and people still report getting Ivy Bridge and Haswell cores sometimes, in 2023; and the better comparison is probably a t2-medium @ $33/mo; a 5-6x markup.
This isn't new information; lambda is crazy expensive; blah blah blah; but I don't hear that dimension brought up enough. Calling back to my previous point: Is a function executing in 75ms instead of 80ms worth paying 30% more? Well, we're already paying 550% more; the fact that it doesn't execute in 75ms by default is abhorrent. Put another way: if Lambda, and other serverless systems like it such as hosted CI runners, enables cloud providers to keep old hardware around far longer than performance improvements say it should be; the markup should not be 500%. We're doing Amazon a favor by using Lambda.
When I moved it, I didn’t need to make any code changes :) I just made a systemd file and deployed it.
If you were comparing e.g. the E5-2667v2 to the Xeon Gold 6334 you would be right, because they have the same number of cores and the 6334 has a higher rather than lower clock speed.
But the newer CPUs support more cores per socket. The E5-2643v2 has 6, the Xeon Platinum 8352V has 36.
To make that fit in the power budget, it has a lower base clock, which eats a huge chunk out of Ice Lake's IPC advantage. Then the newer CPU has around twice as much L3 cache, 54MB vs. 25MB, but that's for six times as many cores. You get 1.5MB/core instead of >4MB/core. It has just over three times the memory bandwidth (8xDDR4-2933 vs. 4xDDR3-1866), but again six times as many cores, so around half as much per core. It can easily be slower despite being newer, even when you're compute bound.
> We've got a project that can build locally, M2 Pro, docker pull and push included, in something like 40 seconds; the CI takes 4 minutes. Its the crusty CPUs; its slow networking; its the "step 1 is finished, wait 10 seconds for the orchestrator to realize it and start step 2".
Inefficient code and slow hardware are two different things. You can have the fastest machine in the world that finishes step 1 in 4ms and still be waiting 10 full seconds if the system is using a timer.
But they're operating in a competitive market. If you want a faster system, patronize a company that provides one. Just don't be surprised if it costs more.
Of course the types of optimisations that a compiler may (or may not) do on aarch64 vs x86_64 are completely different and may explain the difference (we actually compile with -march=haswell for x86_64), but generally Graviton seems like a really good deal.
Edit: yes, haswell:-)
It's not useful at all. It effectively measures IPC (instructions per clock), which is just chip vendor bragging rights.
Assuming that all the chips meet some baseline performance criteria: for datacenter and portable devices, the real benchmark would be "instructions per joule."
For desktop devices "instructions per dollar" would be most relevant.
For cloud customers as well
Cloud costs are dominated by power delivery and cooling. Both of those are directly influenced by how much power the chip uses to achieve it's performance target.
I guess it does indirectly influence dollar cost, but I was referring to MSRP of the chip. As a simple example: the per-chip cost of Graviton is probably enormous (if you factor R&D into the cost of a chip), but it's still cheaper for Amazon customers. Why? Power and cooling.
I don't understand where these power and cooling mantras came from, but large-scale cloud providers have very low PUE (Google publishes historical data at https://www.google.com/about/datacenters/efficiency/). That means you can take the basic power from a thread, add some for memory, and multiply by just a bit to get Watts. Take that and plug in your favorite $/kWh guess and get a price.
Ignoring Graviton, which doesn't have published power data to my knowledge, you can look up the TDP for a bunch of chips and see that it's a a few Watts per thread [1]. Similar calculations can be done for RAM. You end up at 5ish Watts per thread with some RAM attached. Let's call it 10 all in with cooling or other stuff. Since 10W is 1% of a kW, we end up with .01 kWh per hour. The top hit for "us power commercial rates" [2] says that we should assume 7c per kWh or so. That means our instance with cooling and overheads, needs to include .01 x .07 => $.0007/hr of power and cooling costs.
A single core w/ 4 GiB of memory on GCP at 3yr commitment rates (so we're focused on the long-term depreciation price) is .009815 + 4x.001316 => $.015/hr or about 20x as much as the power.
tl;dr: Power costs add up, but they are not even close to dominating the costs of cloud pricing.
[1] https://wccftech.com/amd-epyc-7h12-cpu-64-core-zen-2-280w-td...
[2] https://www.statista.com/statistics/190680/us-industrial-con...
I think the misconceptions come from enterprise DC environments with traditional hot/cold aisle designs, servers running way cooler than they need to be and peak power requirements leading to overly expensive power/cooling costs. PUE for these closer to 2.0 than GCE is to 1.0.
If you have something like GCE where you can control for all of those, i.e run the DCs hot, eliminate transient peaks, source power cheaply and use super efficient cooling like evaporative or geothermal pumped water etc then yeah, it's a completely different ballgame.
It's pretty safe to assume all the hyperscalers are doing all of these things too, they aren't stupid. :)
+1. Moreover the author then seems to conclude from this benchmark:
> Overall, these numbers suggest that the Qualcomm processor is competitive.
This is an odd conclusion to draw from this test and these numbers, given how little this benchmark tests (just string operations). Does this benchmark want to test raw CPU power? Then why "normalize" to 3GHz? Does it want to test CPU capabilities? If so why use such a "narrow" test?
IMO this benchmarking does for a good data-point, but far from enough to draw much of a conclusion from.
Well, it's one measure, but scaling potential and absolute performance limit matters as well. In use cases where a desktop is performing work to help one or more humans, the cost of salary and other support may utterly drown out even a very expensive extremely high power system. Ie., a 1200W workstation being used at maximum power 10 hours a day 365 days a year (an absurd utilization ratio) at $0.15/kWh (average US electricity cost) would still only be around $650/year. If it boosted the productivity of a typical tech worker even 1% it'd pay for itself no problem. It's easy to lose sight sometimes of how historically incredible the bang for the buck is in tech.
I think it's important to remember because some designs that do incredibly well at small sizes run into challenges scaling. Like with CPUs, the nature of silicon fabrication makes it ever more difficult to grow a monolithic die. Switching to a chiplet-based design does mean absolute efficiency and minimum power challenges amongst others, but dramatically improves scalability at the high end. It's a sort of infrastructure tradeoff. Apple for example has done incredibly well with its big silicon SoCs from handhelds to portable Macs, but it's been struggling to do a Mac Pro or even updated Studio. It's not clear that they physically can do something on the level of a modern Epyc chip anymore than Intel could with monolithic Xeons.
So normalized performance/watt and performance/$ and so on do matter, but absolute final density and scale up are another part of the matrix for certain use cases too.
That aside, I think it would be more important to look at the carbon emissions impact of code/cpu's.
That one person parsing urls might save few hundred dollars of not optimising their code but if it is being used billions of times a day the energy (and therefore emissions) impact can be huge.
I’m not trying to optimize for cost or energy efficiency.
DSPs will beat CPUs on DSP workloads, but, as expected, utterly fail for any general purpose workload.
so I tossed it on my 5950x and it reports "ns/url=148.994" which is about right from what I understand (zen4 gets a nice ~30% bump on a number of Linux benchmarks over zen3), here its 35% with only a 16% clock advantage.
Which is about the diff between the M2 and my old 5950X in a generation older process/etc.
So, yah current M2 slower than previous gen Intel and AMD both in this benchmark.
Density is important because if you can only have a max of, for example, 20 cores for an ARM solution, but 96 cores in the case of EPYC Genoa, Genoa is going to win out for any multi-core workload.
The strength of Apple silicon is that it can crush benchmarks and transfer that power very well to real world concurrent workloads too e.g. is this basically just measuring the L1 latency? If not are the compilers generating the right instructions etc. (One would assume they are but I have had issues with getting good arm codegen previously, only to find that the compiler couldn't work out what ISA to target other than a conservative guess)
Wouldn't modern C++ compilers have decent codegen tuning for all these platforms?
If the benchmark were only measuring L1 latency, what would that imply about the ‘scaling by inverse clock speed’ bit? My guess is as follows. Chips with higher clock rates will be penalised: (a) it is harder to decrease latencies (memory, pipeline length, etc) in absolute terms than run at a higher clock speed to maybe do non-memory things faster; and (b) if you’re waiting 5ns to read some data, that hurts you more after the scaling if your clock speed is higher. The fact that the M1 wins after the scaling despite the higher clock rate suggests to me that either they have a big advantage on memory latency or there’s some non-memory-latency advantage in scheduling or branch prediction that leads to more useful instructions being retired per cycle.
But maybe I’m interpreting it the wrong way.
The 8cx 3rd Gen is based on Cortex X1, the same as Snapdragon 888.
The Graviton 3 is based on Neoverse V1 which in itself is tweaked version of Neoverse N1 with relaxed pref / watt / die area and much improved SIMD work load. And N1 is based on Cortex X1.
The current Snapdragon is on Cortex X3. With Cortex X4 coming soon.
- on an AMD Ryzen 5 Pro 4650U (laptop) I get 270 ns/url
- on an Ampere Altra Max M128 I get 340ns/url
So yeah, this show how great Apple's M2 are
We're discussing overall compute power differences between CPU architectures, minute differences in performance between identical-architecture CPU cores stemming from higher clock speeds is outside the scope of this discussion.
Apple uses LPDDR4 modules soldered to a PCB, sourced from the same Korean company that everyone else uses. Intel has used the exact same architecture since Cannon Lake, in 2018.
Last I checked apple M1 max chips have up to 800GB/s throughput, whilst AMDs high end chips taper out at around ~250GB/s or so, closer to what a standard M2 chip does (not max, or pro version). at the top end they've got at least 2x the memory bandwidth than other CPU vendors, and that's likely the case further down too.
(The M1 Ultra is the one with 800GB/sec bandwidth on paper)
Where it gets interesting is how access to the memory is multiplexed in the most extreme case where 3x CPU clusters, 32x GPU and 16x ANE cores are attempting to fetch memory blocks at different locations and at once. It is not unreasonable to presuppose that Apple Silicon contraptions use the switched memory architecture, however with such a high degree of parallelism it is very intriguing to know how the memory architecture has been actually designed and optimised. High performant memory access has always been a big deal that usually comes with big money attached to it via a, naturally, non-memory bus associated connection.
What is somewhat relevant as the source of confusion here is that Apple puts the DRAM on the same package as the processor rather than on the motherboard nearby like is almost always done for x86 systems that use LPDDR. (But there's at least one upcoming Intel system that's been announced as putting the processor and LPDDR on a shared module that is itself then soldered to the motherboard.) That packaging detail probably doesn't matter much for the entry-level Apple chips that use the same memory bus width as x86 processors, but may be more important for the high-end parts with GPU-like wide memory busses.
> Apple uses LPDDR4 modules soldered to a PCB, sourced from the same Korean company that everyone else uses.
Apple might have sourced and soldered on LPDDR modules from the same company, but it is not LPDDR4 and it is LPDDR5 6400 connected via a 256 (Pro), 512 (Max) or 1024 (Ultra) bit wide memory bus.
> Intel has used the exact same architecture since Cannon Lake, in 2018.
Other than LPDDR5 6400 had not existed in 2018, no Intel CPU has ever used a 512, leave alone 1024, bit wide memory bus even in the server setup. Wide memory bus is conducive of faster concurrent complex builds, and Rust / Haskell builds show significantly faster build speeds as well.
> Rust, cargo build for our "test world binary".: 64 core Threadripper 3990x @ PBO level 3, ~400w, with Optane drive - 1 min 37s M1 Max, 64 GB, on battery @ 30% charge - 1:34s
[0]: https://github.com/linux-surface/surface-pro-x/issues/43
They're also not the same core architecture? Comparing ARM chips that conform to the same spec won't necessarily scale the same across frequencies. Even if all of these CPUs did have scaling clock speeds, their core logic is not the same. Hell, even the Firestorm and Icestorm cores on the M1 SOC shouldn't be considered directly comparable if you scale the clock speeds.
It seems with Apple now focusing on AARCH64 the architecture being taken seriously by developers (because ya know, phones, raspberry pis etc weren't enough :)) and there is good support for most software I use day-to-day.
Now it seems Amazon is pushing HARD on Graviton, as the release of 3 has seen many services now being featured running on it. They made great process with Graviton 3, so it's not hard to see why.
Exciting times :)
Not to mention performance and energy consumption benefits.
The phone had it to begin with due to the low consumption and easy instructions set. Then the Moore's law did the rest and now they're powerful enough for data centers
Next year it’s likely all iPhones (plus possibly iPads) will also be on the new process.
It looks like Apple sells around 200 million iPhones a year, and the two pro models are somewhat more popular combined than the non-pros.
So even if we assume 2/3rds of Apple sales are older models, the pros would still need around 40 million chips on the new process in the first year.
For comparison it looks like AMD sells about 80 million chips a year, across all CPU models.
https://wccftech.com/tsmc-faces-order-cutback-from-major-5nm...
Ampere A1 Compute instances