Also, I'm not sure if "correcting" the numbers for 3 GHz is reasonable and reflects real-life performance. Perhaps some throttling could be applied to test the CPUs using a common frequency?
Also, I'm not sure if "correcting" the numbers for 3 GHz is reasonable and reflects real-life performance. Perhaps some throttling could be applied to test the CPUs using a common frequency?
It's not useful at all. It effectively measures IPC (instructions per clock), which is just chip vendor bragging rights.
Assuming that all the chips meet some baseline performance criteria: for datacenter and portable devices, the real benchmark would be "instructions per joule."
For desktop devices "instructions per dollar" would be most relevant.
+1. Moreover the author then seems to conclude from this benchmark:
> Overall, these numbers suggest that the Qualcomm processor is competitive.
This is an odd conclusion to draw from this test and these numbers, given how little this benchmark tests (just string operations). Does this benchmark want to test raw CPU power? Then why "normalize" to 3GHz? Does it want to test CPU capabilities? If so why use such a "narrow" test?
IMO this benchmarking does for a good data-point, but far from enough to draw much of a conclusion from.
Well, it's one measure, but scaling potential and absolute performance limit matters as well. In use cases where a desktop is performing work to help one or more humans, the cost of salary and other support may utterly drown out even a very expensive extremely high power system. Ie., a 1200W workstation being used at maximum power 10 hours a day 365 days a year (an absurd utilization ratio) at $0.15/kWh (average US electricity cost) would still only be around $650/year. If it boosted the productivity of a typical tech worker even 1% it'd pay for itself no problem. It's easy to lose sight sometimes of how historically incredible the bang for the buck is in tech.
I think it's important to remember because some designs that do incredibly well at small sizes run into challenges scaling. Like with CPUs, the nature of silicon fabrication makes it ever more difficult to grow a monolithic die. Switching to a chiplet-based design does mean absolute efficiency and minimum power challenges amongst others, but dramatically improves scalability at the high end. It's a sort of infrastructure tradeoff. Apple for example has done incredibly well with its big silicon SoCs from handhelds to portable Macs, but it's been struggling to do a Mac Pro or even updated Studio. It's not clear that they physically can do something on the level of a modern Epyc chip anymore than Intel could with monolithic Xeons.
So normalized performance/watt and performance/$ and so on do matter, but absolute final density and scale up are another part of the matrix for certain use cases too.
That aside, I think it would be more important to look at the carbon emissions impact of code/cpu's.
That one person parsing urls might save few hundred dollars of not optimising their code but if it is being used billions of times a day the energy (and therefore emissions) impact can be huge.
I’m not trying to optimize for cost or energy efficiency.
For cloud customers as well
Cloud costs are dominated by power delivery and cooling. Both of those are directly influenced by how much power the chip uses to achieve it's performance target.
I guess it does indirectly influence dollar cost, but I was referring to MSRP of the chip. As a simple example: the per-chip cost of Graviton is probably enormous (if you factor R&D into the cost of a chip), but it's still cheaper for Amazon customers. Why? Power and cooling.
I don't understand where these power and cooling mantras came from, but large-scale cloud providers have very low PUE (Google publishes historical data at https://www.google.com/about/datacenters/efficiency/). That means you can take the basic power from a thread, add some for memory, and multiply by just a bit to get Watts. Take that and plug in your favorite $/kWh guess and get a price.
Ignoring Graviton, which doesn't have published power data to my knowledge, you can look up the TDP for a bunch of chips and see that it's a a few Watts per thread [1]. Similar calculations can be done for RAM. You end up at 5ish Watts per thread with some RAM attached. Let's call it 10 all in with cooling or other stuff. Since 10W is 1% of a kW, we end up with .01 kWh per hour. The top hit for "us power commercial rates" [2] says that we should assume 7c per kWh or so. That means our instance with cooling and overheads, needs to include .01 x .07 => $.0007/hr of power and cooling costs.
A single core w/ 4 GiB of memory on GCP at 3yr commitment rates (so we're focused on the long-term depreciation price) is .009815 + 4x.001316 => $.015/hr or about 20x as much as the power.
tl;dr: Power costs add up, but they are not even close to dominating the costs of cloud pricing.
[1] https://wccftech.com/amd-epyc-7h12-cpu-64-core-zen-2-280w-td...
[2] https://www.statista.com/statistics/190680/us-industrial-con...
I think the misconceptions come from enterprise DC environments with traditional hot/cold aisle designs, servers running way cooler than they need to be and peak power requirements leading to overly expensive power/cooling costs. PUE for these closer to 2.0 than GCE is to 1.0.
If you have something like GCE where you can control for all of those, i.e run the DCs hot, eliminate transient peaks, source power cheaply and use super efficient cooling like evaporative or geothermal pumped water etc then yeah, it's a completely different ballgame.
It's pretty safe to assume all the hyperscalers are doing all of these things too, they aren't stupid. :)
DSPs will beat CPUs on DSP workloads, but, as expected, utterly fail for any general purpose workload.
so I tossed it on my 5950x and it reports "ns/url=148.994" which is about right from what I understand (zen4 gets a nice ~30% bump on a number of Linux benchmarks over zen3), here its 35% with only a 16% clock advantage.
Which is about the diff between the M2 and my old 5950X in a generation older process/etc.
So, yah current M2 slower than previous gen Intel and AMD both in this benchmark.
Of course the types of optimisations that a compiler may (or may not) do on aarch64 vs x86_64 are completely different and may explain the difference (we actually compile with -march=haswell for x86_64), but generally Graviton seems like a really good deal.
Edit: yes, haswell:-)
https://buildjet.com/for-github-actions/blog/a-performance-r...
I wish this were only a theoretical concern, a theoretical incentive, but its not. Github Actions is slow, and Gitlab suffers from a similar problem; their hosted SaaS runners are on GCP n1-standard-1 machines. The oldest machine type in GCP's fleet, the n1-standard-1 is powered by a variety of dusty, old CPUs Google Cloud has no other use for, from Sandy Bridge to Skylake. That's a 12 year old CPU.
But the cloud provider prefers the latter because it has 500% more cores for 50% more power. Which is why the latter still goes for >$2000 and the former is <$15.
It really does not depend on the workload, when those workloads we're talking about are by-and-large bounded to 1vCPU or less (CI jobs, serverless functions, etc). Ice Lake cores are substantially faster than Ivy Bridge; the 8352V will be faster in practically any workload we're talking about.
However, I do agree with this take, if we're talking about, say, lambda functions. The reason being that the vast majority of workloads built on lambda functions are bounded by IO, not compute; so newer core designs won't result in a meaningful improvement in function execution. Put another way: Is a function executing in 75ms instead of 80ms worth paying 30% more? (I made these numbers up, but its the illustration that matters).
CI is a different story. CI runs are only bound by IO for the smallest of projects; downloading that 800mb node:18 base docker image takes some time, but it can very easily and quickly be dwarfed by all the things that happen afterward. This is not an uncontroversial opinion; "the CI is slow" is such a meme of a problem at engineering companies nowadays that you'd think more people would have the sense to look at the common denominator (the CI hosts suck) and not blame themselves (though, often there's blame to go around). We've got a project that can build locally, M2 Pro, docker pull and push included, in something like 40 seconds; the CI takes 4 minutes. Its the crusty CPUs; its slow networking; its the "step 1 is finished, wait 10 seconds for the orchestrator to realize it and start step 2".
And I think we, the community, need to be more vocal about this when speaking on platforms that charge by the minute. They are clearly incentivized to leave it shitty. It should even surface in discussions about, for example, the markup of lambda versus EC2. A 4096mb lambda function would cost $172/mo if ran 24/7, back-to-back. A comparable c6i-large: $62/mo; a third the price. That's bad enough on the surface, and we need to be cognizant that its even worse than it initially appears because Amazon runs Lambda on whatever they have collecting dust in the closet, and people still report getting Ivy Bridge and Haswell cores sometimes, in 2023; and the better comparison is probably a t2-medium @ $33/mo; a 5-6x markup.
This isn't new information; lambda is crazy expensive; blah blah blah; but I don't hear that dimension brought up enough. Calling back to my previous point: Is a function executing in 75ms instead of 80ms worth paying 30% more? Well, we're already paying 550% more; the fact that it doesn't execute in 75ms by default is abhorrent. Put another way: if Lambda, and other serverless systems like it such as hosted CI runners, enables cloud providers to keep old hardware around far longer than performance improvements say it should be; the markup should not be 500%. We're doing Amazon a favor by using Lambda.
If you were comparing e.g. the E5-2667v2 to the Xeon Gold 6334 you would be right, because they have the same number of cores and the 6334 has a higher rather than lower clock speed.
But the newer CPUs support more cores per socket. The E5-2643v2 has 6, the Xeon Platinum 8352V has 36.
To make that fit in the power budget, it has a lower base clock, which eats a huge chunk out of Ice Lake's IPC advantage. Then the newer CPU has around twice as much L3 cache, 54MB vs. 25MB, but that's for six times as many cores. You get 1.5MB/core instead of >4MB/core. It has just over three times the memory bandwidth (8xDDR4-2933 vs. 4xDDR3-1866), but again six times as many cores, so around half as much per core. It can easily be slower despite being newer, even when you're compute bound.
> We've got a project that can build locally, M2 Pro, docker pull and push included, in something like 40 seconds; the CI takes 4 minutes. Its the crusty CPUs; its slow networking; its the "step 1 is finished, wait 10 seconds for the orchestrator to realize it and start step 2".
Inefficient code and slow hardware are two different things. You can have the fastest machine in the world that finishes step 1 in 4ms and still be waiting 10 full seconds if the system is using a timer.
But they're operating in a competitive market. If you want a faster system, patronize a company that provides one. Just don't be surprised if it costs more.
When I moved it, I didn’t need to make any code changes :) I just made a systemd file and deployed it.
Density is important because if you can only have a max of, for example, 20 cores for an ARM solution, but 96 cores in the case of EPYC Genoa, Genoa is going to win out for any multi-core workload.