In all other respects the M1 stomps the TGL.
In all other respects the M1 stomps the TGL.
Traditional POSIX mutexes in MacOS weren’t heavily optimized. When I worked at Apple ~5 years ago they still didn’t have a good futex implementation (not sure if that’s changed since). Generally the OS team doesn’t care because you should be using libdispatch.
Anyway I’d love to read up on where you encountered this claim so I can better understand the countours of how this thing works.
go test -test.bench=Cond2 --test.run=ZZ --test.cpu=1,2,4 sync
For me: goos: linux
goarch: amd64
pkg: sync
BenchmarkCond2 3575089 336 ns/op
BenchmarkCond2-2 3236198 370 ns/op
BenchmarkCond2-4 2831134 420 ns/op
goos: darwin
goarch: arm64
pkg: sync
BenchmarkCond2 5903233 190.4 ns/op
BenchmarkCond2-2 85038 27434 ns/op
BenchmarkCond2-4 2500750 677.9 ns/op
Way, way worse in the 2- and 4-thread cases on the mac. I don't know why, and the performance of the 2- case swings all over the place.There was another comment here that was deleted that accidentally showed the output of running x86-64 go on the M1. That's another weird risk at the user level, that you may unknowingly run the wrong binary through rosetta2 instead of running the native one.
goos: darwin
goarch: arm64
pkg: sync
BenchmarkCond2 4849542 249.7 ns/op
BenchmarkCond2-2 2639139 531.1 ns/op
BenchmarkCond2-4 3112030 393.1 ns/ophttps://github.com/golang/go/commit/c15593197453b8bf90fc3a90...
https://github.com/chriselrod/VectorizationBase.jl/issues/22
I do a lot of SIMD optimization, so it'll be interesting to see how the M1 competes once there's native support in Julia.