How does macOS manage virtual cores on Apple Silicon?
eclecticlight.co
eclecticlight.co
By the way, I was confused by how you equated architecture and RTL. In this context I use "architecture" to mean ISA. I use "microarchitecture" to refer to how e.g. its performance characteristics are realized, like IPC.
You seem to know more than I do and would love you to comment. Thanks!
It was designed for a large number of equal cores (aka SMP), meaning people did dispatch_async to the default concurrent queues all over the place, which is a bad pattern when you have to shrink down to phone size. Also, dispatch_semaphore has priority inversions and a lot of other features (like dispatch_group) are semaphores in disguise.
It does work if you use it carefully, but Swift concurrency is a different design for a reason.
https://developer.apple.com/documentation/apple-silicon/tuni...
"On a Mac with Apple silicon, GCD takes into account the differences in core types, distributing tasks to the appropriate core type to get the needed performance or efficiency."
Wait really? Where can I read more about this, this goes against what I would assume.
Edit: Or do you mean that when dispatch_async the block is run with the QoS of target queue instead of source? That is what I would normally expect, if you want to "inherit priority" then dispatch_sync would do that, at the expense of blocking.
Tangent: With today's insanely powerful hardware you should not ever be constrained in your programs to ever have to consider setting core affinity and if you go down that road you might want to reevaluate what you're doing, because you're probably doing something wrong and blaming it on the OS scheduler. Even constrained to the E cores (which are still plenty fast), your programs should perform well. I think more developers need to start writing software on slow machines on purpose because too many apps are written on machines that cost $4000+ with the newest chip and gpu innovations and then are never tested on slower more commonly used hardware, and things end up being dog slow on those machines and get no attention. If more apps were built on slow hardware and were still fast, they'd be even more so on the $4k machines. The macbook air was great for this because it was fanless and every core was essentially an "E" core, forcing you to optimize the code you wrote for the selfish reason of it not being annoying to run the code you were writing. Even if selfish, the net result was production code that was blazing fast once deployed on server grade hardware.
That is wrong in HPC environment.
Most programs tailored for HPC will set core affinity manually. And they have very good reason to do so (cache affinity, memory bindings).
The correct abstraction depend of your domain. There were indeed very little incentive up to know to play with core affinity and the scheduler in a classical desktop environment.
The story is different on systems with strong constraints on efficiency.
https://en.wikipedia.org/wiki/Asymmetric_multiprocessing
IIRC big.LITTLE implementations tended to have cores that didn't support the same instruction sets, meaning you couldn't migrate tasks between them if you needed to. Kind of like how laptops could switch between integrated and discrete GPUs, but some users would need to switch to the discrete GPU to use an external monitor even if they didn't want the power hit.
Also, "big.LITTLE" is a pretty strange brand name.
https://en.wikipedia.org/wiki/ARM_big.LITTLE#Heterogeneous_m...
https://en.wikipedia.org/wiki/Heterogeneous_computing
It even lists big.LITTLE there as a typical example in the second article. big.LITTLE itself never had different ISAs as far as I could tell, just scheduling caveats that lead to efficiency tradeoffs, like the first article mentions.
This let them run the same code, even the same system level code. The scheduler only had to optimize performance, without tasks being pinned to one core or another for correctness.
The industry didn’t have a choice. The market was demanding higher performance within the same thermal envelope and the same energy consumption. These are mobile devices, you can’t put a bigger heat sink on it and then crank up the power.
You can find tons of academic literature discussing the necessity of this development (along with many other things that have come to pass), how it would work, etc. in the decade leading up to the introduction. ARM didn’t just release it to the world and say “Surprise!”
We knew it was coming, we just didn’t do the best job of preparing for it.
Much smaller silicon and software changes would have been needed to allow for just two different clock rates.
What was the argument for designing two different types of cores instead?
The efficiency cores are physically smaller, cheaper, etc. One you reduce the performance expectations to the point where they satisfy the requirements, they are a much better choice. The efficiency core has more perf-per-watt than a clocked-down performance core.
It seems counterintuitive. But when your budget for transistors is X and your budget for power draw/heat dissipation is Y, and there is no leeway whatsoever, the big.LITTLE concept gets you more aggregate performance.
From what I understand the transistors aren't mostly for enabling the higher clock speed at all, they allow for wider cores that do more per cycle. It doesn't seem clear at all that they would be less efficient than a narrower design that does less per cycle, if clocked much lower to yield the same final performance per cycle.
Not only are the functions commented out, there's this note
THREAD_AFFINITY_POLICY:
This policy is experimental.
This may be used to express affinity relationships
between threads in the task. Threads with the same affinity tag will
be scheduled to share an L2 cache if possible. That is, affinity tags
are a hint to the scheduler for thread placement.
The namespace of affinity tags is generally local to one task. However,
a child task created after the assignment of affinity tags by its parent
will share that namespace. In particular, a family of forked processes
may be created with a shared affinity namespace.
https://github.com/apple-oss-distributions/xnu/blob/main/osf...> Use it to automatically run busy background apps on the M1 or M2's efficiency cores to save power, leaving the performance cores for the apps you want to run fastest.
As far as I know, it allocates background/low priority tasks to little cores such as OS housekeeping tasks. When you're on low powered mode or low battery mode, it automatically switches to little cores to save energy. When you need extra multithread power, it uses all cores.
That's why chromebooks are an idea that was perhaps too early (and hardware too anaemic)
> App Tamer can take special advantage of Apple Silicon powered Macs, which have two different types of processor cores. Use it to automatically run busy background apps on the M1 or M2's efficiency cores to save power, leaving the performance cores for the apps you want to run fastest.
- shouldn't that be 'P cores'?
> it’s apparent that during Game Mode, the game was given exclusive use of the two E cores, and threads from other processes fixed at low QoS, which would require them to be run on the E cores, were kept waiting. The game’s threads were run on a combination of E and P cores, with much of their load being concentrated on the E cores. This appears to be energy-efficient, and ideal for use on notebooks running on battery power.
So: the E cores are resevered for the use of games, but the game still makes use of some P cores.
https://eclecticlight.co/2023/10/18/how-game-mode-manages-cp...
I think there are a lot of otherwise graphically simple games that poll or sit there in busy wait loops or render more frames than the display can show. It’s like they’ve been built for a resource starved environment and don’t know how to properly handle abundance. It’s really frustrating!
You should at least make your frame limit be equal to your screen refresh rate.
GSync isn't magic. It can only do so much. In this situation, where you've tied your frames to 24 (for some reason?) and your monitor refresh rate is 165Hz, GSync isn't going to save you. GSync will help you if there are sudden, brief drops in FPS below the refresh rate and keep things in sync. It will not save you if your refresh rate is 165Hz and your frame rate is 24...
> Maybe when electricity prices go back down, I'll turn it up
I don't think you'll notice the extra $5 a month it costs to run that a couple hours a week.
We should at least use a real reason why you'd do this. Saving a smidge of power is not one of them.
You should, instead, under-clock your monitor refresh rate to something low, like 75 or 60Hz, then frame-limit to that number as well.
Downclock your monitor to something like 60Hz, then pin your frames there too.
0.2 kw * 8 hours * 30 days * €0.60 =
1.6 kw/h/d * 30 days * €0.60 =
48 kw/h * €0.60 =
€28.80 per month
I game about 3-6 hours a day, but I also do a lot of (unreal) programming on the same computer. So according to measurements, it draws about 2-3kw/h a day. Without limiting the framerate, it's closer to 10kw/h per day. We're talking a savings of ~€4 per day. That's over €100 per month in savings.
> Downclock your monitor to something like 60Hz, then pin your frames there too.
I'm not sure what my monitor rate has to do with fps. They are two totally independent measurements that are very rarely synced if you are doing any kinds of graphically intensive work, even if you pin them to the same rate. The monitor will still double-post frames every so often, or even skip a frame. Another good example is a paused video which is 0fps, but the monitor doesn't care. It just keeps showing the same frame over and over again. The same thing happens here, and with (g|v)sync, there's never any tearing.
Again, your setup is so far from ideal, you should reconsider. Reduce your refresh rate to some multiple of 24 if you insist on 24 for some unknown reason.
It's one of those situations where you are being so clever you're hurting yourself and not realizing it.
They picked a random 24 FPS target (maybe based off old movies for some reason?), but their monitor refreshes at 165Hz. The mis-match of frames and refreshes is so off, they will experience tearing even with mostly-static backgrounds.
Basically, nothing the parent poster is doing makes sense. Neither for power savings nor viewing pleasure.
Not random, and not "old movies". Modern day movies are still 24fps, it's part of what makes movies look the way they do, motion blur.
> but their monitor refreshes at 165Hz
So? It refreshes the image 165 times a second, regardless of video framerate. It just so happens that many of those 165 refreshes in this case will be the same image over and over. Pause a video, so 0 fps, or 1 fps depending on your take, the monitor is still refreshing 165 times a second because it doesn't have anything to do with the video in this case.
You will experience tearing with this much of mis-alignment. It's why things like GSync/FreeSync even exist.
> motion blur
You can enable motion blur in most game settings...
If the game is a rpg or any casual genre where timing is not important and lack of it doesn’t impact gameplay or competitiveness, what’s it matter if they run at 24fps? Maybe they run that low to screen capture at a lower frame rate to not needlessly burn cpu converting 60 or 160 fps down to 24? And most games cap at 30 or 60, yet higher refresh rate monitors don’t have tearing issues or everyone would be throwing a fit. My son has a 120hz monitor and plays countless games capped at 60. Is it the fact they all divide into the refresh rate that stops the issue?
It happens in both directions.
Yes, staying at some multiple of your refresh rate is the general advice. So 60 would be ok on a 120Hz screen.
The OP's comment that started all this, is doing 24fps at 165hz, which is 6.875. They claim to be in the gaming world too, which makes their bizarre decision even more puzzling. Just don't do this.
Even if it is exactly a multiple or even the same number, without vsync, you’ll still get tearing. Why? Because fps is a dynamic number. It’s a measure of how many frames your card can generate in a second. By the time you measure, the frames are already output and gone from the buffer.
The refresh rate on the monitor is a constant, never changing value.
Even if you set the fps to 30, you’ll sometimes render 30.1 or 29.9 frames in a second due to random jitter from the geometry/shader complexity quickly changing on the screen.
There is no such thing as constant fps.
This is better, because we're talking about the lower range right now anyway.
That said, I think the difference between 144Hz and 72Hz is a lot more subtle and can't be seen if you try to glance at them directly. Try viewing them with peripheral vision and there's a difference in smoothness. That shows up in things like moving a window around (or even just moving a mouse). It makes a huge difference in input latency/response, but the difference can't be seen easily by just staring directly at a moving image.
I went outside and took a 4k video at 24fps and played it back. It looks smooth as butter, just like a movie. Then I took the same video, but recorded at 60fps and re-encoded it in 24fps. It looks like the example on this website, where it looks jumpy.
I assume this is because of motion blur.
20fps uses 70 watts, full speed is 200 watts. Times 5 hours, times 30 days, times €0.60.
i still have the antiquated view that performance and efficiency are essentially the same metric and these are the only kinds of ideas i can come up with that would separate them..
So they simply can’t process as much clock for clock.
Deeper speculation, larger out of order window, more accurate branch prediction, larger caches, more execution pipes, larger register files…
That... sucks
It's there to be exposed if the developer thinks it should. That seems like the right balance to me.
I seem to remember years ago I could open activity monitor and watch processes migrate back-and-forth between cores for seemingly no reason instead of just sticking in places.
I haven’t watched to see what happens recently. Is that still an issue?
gamers have been binding to single processors/cores for years for that reason, it's a huge cause of microstutter type behaviors -- but who knows if its perceivable with all the magic going on with processors nowadays.
When a process gets descheduled and then later rescheduled, there is no particular reason to reschedule it on the same core it was on previously. Desktop class hardware doesn't have NUMA. Last level caches might have locality-dependent performance characteristics, but this is usually pretty nuanced, wildly variable between hardware microarchitectures, and not accounted for / not worth accounting for in the OS's scheduler, since there's a good chance that data won't be around next timeslice anyway.
Binding a processor to a single core affinity-wise is not changing any of this.
If you really have a throughput-bound process that needs to squeeze every single microsecond of compute and/or you're super latency bound for some reason, you schedule a process with what's called realtime priority, which changes the scheduling decision of WHEN to deschedule it. Not every operating system has equivalent mechanisms for this, but that's more in line with the phenomena that you're describing. But a process may well cause itself to be descheduled anyway by doing something that puts it to sleep. You have no control over this.
In reality, looking at which core a process is on is mostly the wrong way to be thinking. You're looking at this on a 1-2 second window (the amount of time it takes for Activity Monitor / Task Manager to update its UI). A full second of time is an absolute eternity to a process. Large amounts of work happen on the order of microseconds and quicker. And in that 1-2 second snapshot that Activity Monitor or Task Manager gave you, the process probably spent time on EVERY core -- and you're just seeing wherever it was in the brief moment when Activity Monitor / Task Manager updated its UI.
> gamers have been binding to single processors/cores for years for that reason, it's a huge cause of microstutter type behaviors
It wouldn't be the first time gamers prescribed placebo tweaks because of magical type thinking.
iStatMenus is a similar utility that still exists and is far more popular.
I will say since I use them for work, their ability to keep running despite opening so much crap is amazing. Browser tab wise. Not just limited to a specific silicon either as I use both intel/m2.
If you open Resource Monitor before the freeze, you can usually see what happened on the graphs. If you're fast enough, you can also click "analyze wait chain" to see something that looks like a stack trace. The graphs and trace should make it pretty obvious why an app is frozen. Eg maybe it's trying to save the history sqlite database and your SSD ran out of SLC cache.
edit: I'm gonna get dinged for OT comment again ha
Browsers have been implementing background tab killing recently though, so just because you have a tab doesn't mean the process is actually backing it anymore.
Though, 1500 tabs in one window is beyond the point at which it impacts Safari's responsiveness, but that's completely unrelated to the contents of the tabs and is instead the fault of taking way too long to update every NSView in the tab bar.
I used to be inbox-zero in all gmail tabs a couple of years ago and I will be yet again some day. The main limitation I'm running into is that a lot of useless emails ("We're making some changes to our PayPal legal agreements") have unpredictable subject lines and are sent from the same address that sends useful emails. I already have hundreds of gmail filters that auto-archive useless emails with predictable addresses or subject lines ("Your statement is ready") that I can't unsubscribe from without closing the account. LLMs should solve this problem some day.
I used to make fun of people who hoarded them, but honestly sometimes just keeping that one tab open until you've exhausted its use is worth it.
Of course I also use ctrl+e (QuickTabs) to jump faster with intellij-like fuzzy search. It's kinda like Chrome's built-in ctrl+shift+a, except better.
Windows 98-brain is hard to get rid of...
Also no, history has terrible search. You can't search within pages, search within a single folder/window, etc.
But now there are built in ways to manage tab groups, which I haven’t used yet, but look neat.
As a person who has terrible short-term memory, I'm genuinely curious. I never hoard my tabs past a certain point 10-20 tabs, it actually slows me down in finding what I'm looking for. If I'm working on specific projects and need to save the tabs for later, I just use tab groups in Chrome and associate with a task so when I context switch, I just open up an existing tab group and close my other non-related ones.
https://www.google.com/chrome/tips/#:~:text=You%20can%20grou....
2) As somebody with ADHD, there's an unacceptable level of mental overhead required for organizing things. I simply don't have any spare bandwidth to add organizing tabs into bookmarks and organizing bookmarks or whatever else into my daily functioning as a human being. It has nothing to do with feeling rewarded.
I hope FPGAs take off one day and software can compile hyper-optimized processors that do very specialized tasks.