When an M1 Mac mini is faster than an M1 Pro: contention and core allocation
eclecticlight.co
eclecticlight.co
This is a nice test and good data, but it seems intuitively correct to me that things would work this way. I guess the unintuitive bit is "if you scale the QoS all the way back to minimum, regular M1 outperforms M1 Pro".
It would be interesting to see which one has the lower power usage during these tasks as well, given that regular M1 is a desktop machine (Mini in this case) and M1 Pro is a laptop.
The powermetrics data will have the info necessary to figure this out - it reports the power draw of each core cluster every time it takes a snapshot.
This points out that (relatively logically) that only works if the cores are under active load computational load.
> It would be interesting to see which one has the lower power usage during these tasks as well, given that regular M1 is a desktop machine (Mini in this case) and M1 Pro is a laptop.
The form factor is not really a factor, the M1 is also present in laptops and I don’t think I’ve seen any evidence that the mini’s M1 is tuned in any way.
That said the previous articles on the subject (“1 + 1 = 4” and “how M1 E cores win”) indicated that the E cluster of the M1+ reach 200mW at full residency and frequency (200% at 2GHz), while the E cluster of the M1 tops out around 165mW. Note that this is for the entire cluster of respectively 2 and 4 cores.
Note that it's a bit harder than that. The ecores on M1 and M1 Pro/Max run at around the same clock for an all-core workload, but the clock on regular M1 is halved when it's running a low priority job specifically.
The Pro/Max will ramp their e-cores even further when running high-QoS tasks (spilling over from the P-cores).
The M1 can turbo slightly when only 1/2 e-cores are in use, but under full load the e cores remain a hair under 1GHz.
In Go, it's pretty trivial to do an experiment where you load a bunch of stuff into memory and have a core per thread race through it, basically doing stream, and if you do this the UI will start stuttering prior to going completely non-responsive. Similarly, on the m1 macs, with nothing more than writes to an SMB mount, you can cause the entire system to wedge [this is trivial to reproduce using rclone against a SMB mount]. There's no reason these should ever happen.
What can I say, I was curious if the advertised crazy levels of M1 pro memory bandwidth were genuine or conflating the all-in system-wide not-actually-visible-to-cores bandwidth hypothetically available.
Guess what, the guy with 20 tractors won!
EDIT: I needed to fix the numbers a bit.
Work that should be completed if resources are available should be at utility. This is a good level if you aren't sure what level to pick.
Work the user is waiting on should be user-initiated.
Work required for interactivity (updating the screen, background drawing, etc) should be user-interactive.
Higher QoS work that blocks on lower QoS work can boost the priority of that work, including across process boundaries in some cases like XPC.
All of these details are subject to change; communicate your intent to the system with the proper QoS class for best results.
You'd think so, but software updates are run by signaling some daemon, and its tasks run on E-cores, which means if you are trying to use Xcode and it wants to install something, you may be waiting a very long time.
When correctly implemented, the process you’re waiting on will have its priority boosted appropriately.
So the question is, has someone found a way to force macos/terminal session/command invocation to run only on certain types of cores, for benchmarking purposes? I would imagine, command akin to caffeinate, but that would force the qos or available cores on all child processes would be ideal, but so far I haven't found a way to do that.
So it is possible to get a bit more control, at least on Apple ARM chips, but it’s intended for games. They even mention professional apps needing high performance are able to get the performance they need with the normal Dispatch APIs. There is no locking threads to cores, but you can set thread affinities as hints if you think certain threads should or should not share L2 cache (at least for intel based Macs). https://developer.apple.com/forums/thread/44002
It's been a good run but I think us 2015 MBP owners are going to find it harder to keep them running. Components break and everything is so integrated now it's increasingly difficult to replace the bad bits without swapping major parts.
So now I carry a 2020 MBA and it's been a learning curve since different software is at various stages of AARM compatibility.
The QoS levels are background (lowest), utility, userInitiated, and userInteractive (highest). All but background will run on P-cores.
Does it make any sense to benchmark a process when the QoS was specifically set to "I don't care how long it takes"?
It sounds like if you set low QoS your process will only run on e-cores, is that correct? If so then "obviously" the M1 Pro/Max would be slower, but I don't think this is an obvious consequence from the end user's point of view. Don't get me wrong it's the logical consequence of the SoC configuration and scheduler rules, and I'm not even sure if its the "wrong" choice from a battery life PoV (I'm curious whether it results in lower battery use per-task)
Not that I'm sure if that's what's happening here. But if so, it seems consistent that the lappy would be slower than the desktop.
Thus sadness ensues :D