Go standard library benchmarks – Intel vs. M1
roland.zone
roland.zone
Is this more or less than the expected improvement in benchmarks if this were an Ice Lake or Ryzen 4000 CPU in an equivalent form factor? Are there similar benchmarks comparing e.g. a 2020 Dell XPS 13 to an M1 Macbook Pro?
> A top level takeaway from this data is that for things that rely on pure Go the M1 machine is, in general, a good bit faster than the Intel machine. For code that relies on highly optimized assembly, like a good chunk of the crypto/ code, the Intel machine wins (although I wouldn’t be surprised to see significant gains for the M1 as arm64 assembly implementations get more attention and are further polished).
Basically Go lacks lots of optimized assembly code for ARM64.
https://github.com/golang/go/blob/master/src/crypto/aes/asm_...
In other benchmarks like geekbench, which are not benchmarking golang's libs, the M1 is massively stronger than Intel on e.g. AES-XTS.
https://www.extremetech.com/computing/317304-benchmark-resul...
So hopefully the M1 gets the same hand-optimizing effort that x86 historically got.
Go libs just need to be tweaked to use the correct chip to do the work.
Looks like they did do AES-XTS on the main die because it is required for encrypted storage.
There was no competition for Intel in the x86 market. They attempted to enter mobile which was a huge and growing market and they failed.
This stagnation may have fuelled development elsewhere (RISC-V?) and now that AMD is "back in the game" and ARM is making its way into server hardware and "desktop", there's real competition again.
But these days, not really. Peel back and peer at the CPU and they all look and work much the same. The CISC implementations typically convert to their own simplified instruction subset internally, and the RISC crowd keep adding fatter and fatter CISC-like instructions.
So its not really an x86 instruction set problem. The legacy that Intel (and, a lesser extent, AMD) seem to struggle to overhaul is in their architecture, not the instruction set itself.
This is proven by the M1 doing binary translation of x86 and still being faster than the Intel chips.
That’s the real kick in the nuts for intel isn’t it. I bet apple could have jumped last year or before and been ahead in native performance. But waiting meant they even beat intel at their own game, while massively handicapped.
Basically Intel didn’t need to innovate to make money the last decade. Now it’s come to back bite them on the ass.
Imagine if we were able to run Linux kernel on such CPUs that didn't care about putting in the low power cores or GPU cores and optimised even more for performance. One can only dream so wild.
Some HPC shops would build racks and racks of those.
Not an expert in HPC, but I think you gain the performance from distributing the data across a vast number of cpu instead of having great cpus to begin with.
Also, these results are OK, because it means new macbook performance won't be diminished because of the switch to ARM, but it's not that much of a perf improvement either: keep in mind this benchmark was run against a 2017 i5 CPU. It's not an anemic CPU, but it's not really hard to beat either.
We've seen 50% to 200% improvements over the latest i9 CPUs too...
This is not going to happen, Apple exited that business in 2010 and have shown no signs of rekindling that interest. I agree with you that there could be a lot of potential in Apple Silicon in 1U rackmounts but I am not holding my breath. Apple is 100% focused on consumers, it is not interested at all in datacentres, it doesn't even run OSX in its own datacentres.
M1 is the most exciting thing to happen in tech this week. Or at least I think so :)
Because it was upvoted. Same as most things that go into the front page, aside from YC promos.
I'm still unconvinced and wait for some good old BLAS...
As for the linked microbenchmarks: I guess with 12MB of cache, there's plenty of memory to place the kernel and some other stuff. I wonder, how this plays out here.
A better way might have been to run the Go1 benchmarks instead. That contains a pre-selected set of benchmarks that tests a variety of workloads.