These benchmarks tell you what performance you get with this array of programs, on linux, with these machines. You want to know something about the inherent performance capabilties of each cpu, which is very hard to figure out, since we will never know when/if any of these programs has had a totally optimal solution for each piece of hardware.
Usually the best you can do is just measure actual programs people use.
I found the data compelling--M2 silicon under Linux is generally competitive against contemporary strong x86 laptops, but if you are sensitive about the performance of a particular workload, choose your device carefully.
Fun to think that Apple almost has a monopoly on RISC technical workstations…
Lammps molecular dynamics is by Sandia national labs, they probably have a good ARM implementation. LeelaChessZero is also a SIMD-implementation IIRC. Both are SIMD-accelerated, but LeelaChessZero is faster on M2, but LAMMPS is faster on AMD.
I don't get it at all. Its really hard to see the pattern. Furthermore, its hard for me to imagine that LAMMPS would be much worse on one system over another, given the authors.
---------
Then comes the "in-between" benchmarks. DaCappo is slower on low-power settings for M2, but high-power setting M2 faster than any setting. Very strange behavior, I can't imagine why this would happen.
ZSTD compression is faster on AMD, but decompression is faster on M2. Same codebase, same authors, but different results over two different runs.
It's certainly what Cloudflare found when they were evaluating early ARM server chips.
>the first benchmark would be the popular zlib library. At Cloudflare we use an improved version of the library, optimized for 64-bit Intel processors, and although it is written mostly in C, it does use some Intel specific intrinsics. Comparing this optimized version to the generic zlib library wouldn’t be fair. Not to worry, with little effort I adapted the library to work very well on the ARMv8 architecture, with the use of NEON and CRC32 intrinsics. In the process it is twice as fast as the generic library for some files.
https://blog.cloudflare.com/arm-takes-wing/
SIMD code has to be hand written for both instruction sets before you are comparing Apples to Apples, as it were.
Look at the benchmarks in the post. Some are consumer programs, but others are supercomputer programs with an enormous amount of optimization effort put into them for a variety of platforms.
For instance, in the example above, Cloudflare doubled the speed of the code by taking advantage of ARM specific instructions in the same way they had previously optimized their code for x86 specific instructions.
Of course, using the availible instructions isn't the same thing as putting in a heavy investment of time to make the code as efficient as possible.
As an example, look at x86 AV1 video encoders. The early ones were incredibly slow, but through many iterations of hand written SIMD code, recent versions have been getting much faster on the same hardware.
ARM / Apple M2 here is just 128-bit vectors. AMD / Intel AVX is 256-bit vectors, which is harder to optimize for (!!) due to the increased parallelism.
NEON itself is kinda crappy, missing a few good data-movement instructions (pshufb doesn't exist in ARM world), but IIRC, there are hard-coded swizzles in NEON that reach high levels of performance. So in this case, its simply easier to write pshufb style code in Intel/AMD systems.
Then again, I don't expect that a BLAS library would use pshufb too much. So you'd probably just have a simple dot-product like vector going up-and-down RAM to maximize your matrix-multiplication speeds.
----------
That's what I mean. LAMMPS isn't "easy" but it should fundamentally be a matrix multiplication (which is highly studied, and well optimized on all platforms). Indeed, the best performing BLAS libraries these days are on NVidia, not Intel or ARM for that matter. (Though Intel's AVX512 matrix multiplications are probably next best, they don't apply to this Apple vs AMD comparison).
You can't just assume a problem exists in general. There's some benchmarks in here that are sufficiently well optimized in ARM yet still come out better for the AMD chips. And vice versa for that matter (I'd expect LeelaZero to favor AMD, so I was surprised to see it better on the Apple M2).
---------
Assuming a Russell's Teapot just because there was a teapot from an unrelated benchmark years ago is kind of a bad form of argument. SIMD-compute (AVX, NEON, SVE, etc. etc.) has become more popular today, and there are more NEON_optimized libraries today than even just 5 years ago. And there are plenty of applications in this Phoronix benchmark where we'd reasonably expect good ARM-NEON optimizations.
I'm reminded of the Handbrake people back in 2018 telling users that the current open source implementations of AV1 encoder ran 6,000 times slower than h.264 encode, and that they wouldn't be adding support for AV1 encode anytime soon, as the code base needed time to be optimized.
See here:
https://github.com/HandBrake/HandBrake/issues/457
Just because you run on a particular instruction set doesn't mean you're using that instruction set in an optimal manner.
In my experience Ryzen support on Linux is far from optimal. Have experienced tons of issues on zen+ and zen 2.
And don't forget that the x64 code is slowed down by tons of security mitigations. I don't think most of these are available on M2 yet.
What kinds of issues have you seen?
I've been running a home server with a Ryzen 7 PRO 4750G and so far it's been stellar.
> slowed down by tons of security mitigations. I don't think most of these are available on M2 yet.
Are they needed at all? I thought the security mitigations were mostly for x86-specific architectural issues. Would you mind to elaborate?
Now the existing discovered ones may not, but I'm sure in time there will be ones for Apple Silicon too. Just like how people for the longest time thought AMD was safer than Intel in this regard, or that ARM wasn't vulnerable.
Regarding the other point: run lscpu and compare enabled mitigations on Zen3 and M2. We still need a proper analysis to figure out which ones must be enabled on M2.