How M1 Macs feel faster than Intel models: it’s about QoS
eclecticlight.co
eclecticlight.co
E.g. compiling Rust code is so much faster that it is not even funny. cargo install -f ripgrep takes 22 seconds on a Mac Book Air with M1, same on a hexacore 2020 Dell XPS 17 with 64GB RAM takes 34 seconds.
The comparison is way too close given the ryzen workstation guzzles power, and the macbook is cheaper, portable and lasts 22 hours on a single charge. If Apple can keep their momentum going year over year with CPU improvements, they'll be unstoppable. For now it looks like its not a question of if I'll get one, but when.
6 seconds on the Ryzen, vs 9 seconds on the M1 air.
`time go build -a`, so not very scientific. Could be attributed to the multicore performance of the Ryzen.
Starting applications on the M1 seems to have significant delays too, but I'm not sure if that's a Mac OSX thing. Overall it's very impressive, I just don't see the same lunch eating performance as everyone else.
The battery life and lack of fans is wonderful.
edit: Updated with the arm build on OSX. 16s -> 9 seconds.
A tool I have enjoyed using to make these measurements more accurate is hyperfine[0].
In general, the 3700X beating the M1 should be the expected result... it has double the number of high performance cores, several times as much TDP, and access to way more RAM and faster SSDs.
The fact that the M1 is able to be neck and neck with the 3700X is impressive, and the M1 definitely can achieve some unlikely victories. It'll be interesting to see what happens with the M1X (or M2, whatever they label it).
All this said, there really is no comparison. I don’t even think about battery for the M1, I leave the charger at home, and it’s 100% cool and silent. It’s a giant leap for portable creative computing.
That's a Mac OSX thing. It's verifying the apps.
https://appletoolbox.com/why-is-macos-catalina-verifying-app...
time GOARCH=amd GOOS=linux go build -a
6s for the M1 too!
In addition: let’s talk upgradability or repairability. Oh wait, Apple doesn’t play that game. You’ll get more mileage on the workstation hands-down.
The only win for those those chips I think is battery efficient for a laptop. But, then why not just VNC into a beastmode machine on a netbook and compile remotely? After all, that’s what CI/CD pipeline is for.
The Thinkpad is obviously no powerhouse, but still works great for general desktop use, ie. browsing, email, document editing, music, video (1080p h264 is no problem). The desktop plays GTA V at around 40-50 FPS at 1080p with maximum settings. And this isn't some premium build, it's a pretty standard Asrock motherboard with Kingston ValueRAM and a Samsung SSD.
Decade-old hardware is still perfectly viable today.
Is this how you work?
This is exactly how I've worked for a number of years now, for my home/personal/freelance work. Usually using a Chromebook netbook ssh'ing into my high spec home server. I'd do the same for work, but work usually requires using a work laptop (MacBook).
OTOH, the machine that I'm connecting to has 32c/64t, half a terabyte of RAM and dozens of TB of storage.
All that and my preferred OS (Manjaro/XFCE), which runs on anything, has been more stable than any Mac I've ever owned. Every update to macOS has broken something or changed the UI drastically and in a way I have no control over...
If I ever switch away from desktops, it will be for a Framework laptop or something similar.
Everything in the machine can be upgraded/fixed so it should be good for a while.
I’m not saying this to be snarky. I just want to emphasize that while M1 is great innovation, I put repairability/maintainability and longevity on a higher pedestal than other things. I also highly value many things a computer has to offer: disk, memory, CPU, GPU, etc. I want to be able to interchange those pieces; and I want to have a lot of each category at my disposal. Given this, battery life is not as important as the potential functionality a given machine can provide.
I suspect the number of people, even developers, for whom 16GB memory is plenty probably greatly exceeds the number who need a beast mode Ryzen. But even then, a large proportion of the devs who might need a Build farm on the back end would be doing that anyway so they might as well have an M1 Mac laptop regardless.
Anyway Mac Pro models will come.
For me, SoC power draw when compiling apps is usually around 15 to 18W on a M1 mac mini.
My example app (deadbeef) compiles in about 60 seconds on that mac mini.
On a 2019 i7 16” MBP (six core), it takes about 90 seconds, and draws ~65W for that period.
So… radically more power efficient.
edit: this is the same version of macOS (big sur), xcode, and both building a universal app (intel and ARM).
Is it all CPU though, or is building ARM artefacts less resource intensive than bulding ones for Intel ?
18 core iMac Pro 13339: https://browser.geekbench.com/macs/imac-pro-late-2017-intel-...
M1 Mac mini 7408: https://browser.geekbench.com/macs/mac-mini-late-2020
My exact sentiments. I've been looking for a gateway into Apple and the M1 Air seems like it. It has now become a matter of time and not just a fleeting thought.
I’ve been blown away with games running on M1. If Apple could up their GPU game as well, that’d be really cool.
I don't use Rust, but a quick search returned for example this issue:
And I really care about rust compile times because I spend a lot of small moments waiting for the compiler to run.
I keep hearing this but have not experienced this in person. I usually get about 5 hours of life out of it. My usual open programs are CLion, WebStorm, Firefox (and Little Snitch running in the background).
However, even with not having IDEs open all the time, and switching over from Firefox to Safari, I’m only seeing about 8 hours of battery life (which is still nice compared with my 2013 MBP that has about 30 minutes of battery life).
I would consider getting a warranty replacement. Something is wrong.
For reference, my M1 Air averages exactly 12 hours of screen-on time (yes, I've been keeping track), and the absolute worst battery life I've experienced is 8.5 hours, when I was doing some more intense dev workflows.
Apple seems to have taken the power reduction with the A14 and M1 on TSMC 5nm, not the performance increase.
>The one explanation and theory I have is that Apple might have finally pulled back on their excessive peak power draw at the maximum performance states of the CPUs and GPUs, and thus peak performance wouldn’t have seen such a large jump this generation, but favour more sustainable thermal figures.
https://www.anandtech.com/show/16088/apple-announces-5nm-a14...
It's like buying a car and modifying to take it racing vs buying a race car, the race car was designed to do this.
This is the position that Apple have set up for themselves with their philosophy and process. It would seem that Intel and AMD have to play a very conservative game with compatibility and building a product that increments support for x86 and x64. They can't make some sweeping change because they have to think about Linux and Windows.
Apple own their ecosystem and can move everything at once (to a large degree.) This also gives an opportunity to design how the components should interact. Incompatibility won't be heavily penalized unless really important apps get left behind. The improvements also incentivize app makers to be there since their developer experience will improve.
Talk to people who design chips. The compatibility barely impacts the chip transistor budget these days, and since the underlying CPU isn't running x86 or x64 instructions, it really doesn't impact the CPU design. There may be some intrinsic overhead coming from limitations of the ISA itself, but even there they keep adding new instructions for specialized operations when opportunities allow.
Everyone who develops for one of their platforms is just used to running that treadmill. An ARM transition is just another thing to update your apps for.
To be fair, that is a relatively high-end desktop CPU, but it's also a massive margin for a last-gen model.
Also has a Samsung 980 1TB drive, and a 3070 GPU; the system truly is extremely quick. (Win10+WSL2).
(I'll have to try that compile in the next day or so; will reply to myself here).
For instance Windows on NTFS etc. is notoriously slow at operations that involve a lot of small files compared to Linux or whatever on ext4.
Unless your Dell was also running OSX, you're probably not comparing hardware here.
Finished release [optimized + debuginfo] target(s) in 34.50s
...
cargo install -f ripgrep 137,66s user 13,00s system 435% cpu 34,632 total
For comparison TR1900x (yeah, desktop, power-guzzler, but also several years old), Fedora 34: Finished release [optimized + debuginfo] target(s) in 25.15s
...
real 0m25,271s
user 4m41,553s
sys 0m7,216sHere's r7 4800HS (35w) on linux:
Finished release [optimized + debuginfo] target(s) in 24.98s
real 0m25.151s
user 4m54.987s
sys 0m6.764s
# git clone https://github.com/BurntSushi/ripgrep && cd ripgrep
# cargo clean
# time cargo build
real 0m23,805s
user 1m11,260s
sys 0m3,806s
How is my crappy laptop on par with your M1? :)
You should benchmark against Linux as well.
There's your problem.
Even Microsoft concedes that Windows is legacy technology now.
EDIT: could you measure C compilation. For example:
git clone --depth 1 --branch v5.12.4 "git://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git"
cd linux
make defconfig
time make -j$(getconf _NPROCESSORS_ONLN)
For me it's slightly below 5 minutes.Are you including download time? On my Ryzen 7 3700X, it's done in 17 seconds.
EDIT: I checked on 1.52.1, which is latest stable and it went down to 54 seconds. So that also makes up a significant difference.
I have since FIXED THIS! The solution is to just not let it sleep. I'm using Amphetamine and since then it has been amazingly fast all the time.
I think these app makers might need to start focus grouping some of these app names. Not only is it going to be harder to search for this app, but it'll inevitably also lead to some humorous misunderstandings.
invaderfizz@FIZZ-5950X:~$ hyperfine 'cargo -q install -f ripgrep'
Benchmark #1: cargo -q install -f ripgrep
Time (mean ± σ): 19.190 s ± 0.392 s [User: 294.890 s, System: 17.144 s]
Range (min … max): 18.352 s … 19.803 s 10 runs
The CPU was averaging about 50% load, the dependencies go really fast when they can all parallel compile, but then the larger portions are stuck with single-thread performance.I have two laptops side by side, one is an M1, the other a HP in a slightly higher price point (more memory, bigger SSD). The challenges Intel has in the form factor are obvious — it’s a 1.2Ghz chip that turbos almost 2x as heat allows.
In any dimension that it can do what you need, the Apple wins. Cheaper, faster, cooler, longer battery endurance. The detriments are the things it cannot do — run Windows/Linux or VMs or tasks that need more memory than is currently available.
If someone could figure out a way to get all MacOS apps —including the system — to use the performance cores, perhaps the battery life would be back down to Intel levels?
Other than to prove an utterly useless point - why would you want to even remotely do that?!?
If my time machine backup takes four times as long but my battery lasts longer still why would I care? The overall experience is a HUGE net improvement.
That's the point being glossed over by the majority of commenters and the point the original author is making - benchmarks are interesting, but nothing beats real world experience and in real world experience there are a suite of factors contributing to the M1 and Apples SOC approach spanking the crap out of their competitors.
There is more to life than rAw PoWeR ;)
yah never mind I don't care why these computers are impressive let’s just play macos chess and browse facebook jfc if you’re not interested why did you even join this conversation
I don't know what point would be proved by maxing out the cores other than confirming what seems to be pretty obvious - a hybrid approach has multiple benefits - not just in power efficiency but the user experience as well.
It's not just one aspect of the design choices - but all of them in concert.
In that regard, the moment feels similar to the mid 2000s, when you suddenly saw a huge uptick in Macbooks at technical conferences: software developers seemed to gravitate in droves to the Mac platform (over windows or linux). Over the last 5+ years, that's cooled a bit. But I wonder if the M1 will solidify or grow Apple's (perceived) dominance among developers?
I don't understand how a company of that size kept up a blunder of that magnitude for half a decade, but a half-broken keyboard with an entire missing row of keys was just a nonstarter.
I guess we're both speculating here, but even if the same people were in charge, I think they breathed a sigh of relief when they realized that the personal computing segment had spiced up again. From about 2014 to 2019, Mac revenue basically plateaued in line with the entire laptop market. People were crazy about phones but laptops had hit a wall.
When you have to sell laptops to people whose yesteryear laptops do everything they need, you start adding random bullshit to the product because you have to capture the market's attention somehow. I think this is how we ended up with the touch bar. It's a step backward, but it's flashy and made the product look fresh(er) despite the form factor being identical to what they were selling in 2013.
I don't know which engineers you hang out with, but the only "engineer" I know who uses an M1 Mac does web design. Besides that, the M1 doesn't support most of the software that most engineers are using (unless you're working in a field already focused on ARM), and the fragility of a Macbook isn't really suited for a workshop environment. Among them, Thinkpads still reign supreme (though one guy has a Hackintosh, so maybe you were right all along?)
> But I wonder if the M1 will solidify or grow Apple's (perceived) dominance among developers?
Developers are going to be extremely split on this one. ARM is still in it's infancy right now, especially for desktop applications, so unless your business exclusively relies on MacOS customers, there's not much of a value proposition in compiling a special version of your program for 7% of your users. Because of that, MacOS on ARM has one of the worst package databases in recent memory. Besides, there's a lightning round full of other reasons why MacOS is a pretty poor fit for high-performance computing:
- BSD-styled memory management causes frequent page faults and thrashes memory/swap
- Quartz and other necessary kernel processes consume inordinate amounts of compute
- MacPorts and Brew are both pretty miserable package managers compared to the industry standard options.
- Macs have notoriously terrible hardware compatability
- Abstracting execution makes it harder to package software, harder to run it, and harder to debug it when something goes wrong
...and many more!
In other words, I doubt the M1 will do much to corroborate Apple's perceived superiority among the tech-y crowd. Unless people were really that worried about Twitter pulling up a few dozen milliseconds faster, I fail to see how it's any better of an "engineering" laptop than it's alternatives. If anything, it's GPU is significantly weaker than most other laptops on the market.
All of the things you mentioned are also true of intel macs, which again are wildly popular among web devs of all kinds.
If you can't explain the popularity of those machines in spite of those limitations, I don't see why I should accept those as reasons why apple arm computers won't be popular.
Well, no, it's all totally made up.
I'm talking a 16GB Intel i7 8-thread Macbook vs a 32 thread twin-socket 72GB RTX2080 beast running Ubuntu 20. The Mac crushes it in terms of feel, fit and finish. I haven't tried M1 yet but I bet it'll one-up my current Intel macbook. I'm quite eager to get one.
> Besides that, the M1 doesn't support most of the software that most engineers are using...
??? Other than CUDA, the macbook meets my needs 95% of the time. I'm mostly want for a native Docker networking experience a small minority of the time. I need a responsive GUI head to act as my development environment. All the heavy lift compute is done on servers/on-prem/in-cloud.
> - BSD-styled memory management causes frequent page faults and thrashes memory/swap
Only under super heavy memory demand. I close some browser tabs or IDE panes (I normally have dozens of both).
> - Abstracting execution makes it harder to package software, harder to run it, and harder to debug it when something goes wrong
Almost everything I do is containerized anyways, so this is moot.
I was squarely one of those "why would anyone use mac? It's overpriced, lock in, $typical_nerd_complaints_about_mac" until COVID happened and it became my daily driver. Now I can't go back.
> - MacPorts and Brew are both pretty miserable package managers compared to the industry standard options.
No snark intended - Like what? I'm not exactly blown away by Brew, but it's been generally on par with Apt(get). Aptitude is marginally better. There's not a single package manager that doesn't aggravate me in some way.
> BSD-styled memory management causes frequent page faults and thrashes memory/swap
You're basically pulling this out of nowhere. Not once in six years of using MacOS has this ever happened to me.
> Quartz and other necessary kernel processes consume inordinate amounts of compute
Yes, because Windows is so much better. This is sarcasm. Just pull `services.msc` and take a look at everything that's running.
> MacPorts and Brew are both pretty miserable package managers compared to the industry standard options.
In your opinion, what is industry standard? `apt-get` and `yum`? I have yet to come across a better package manager than Brew. Brew just works. Additionally, most binaries that are installed in Brew don't require elevation. Which is fantastic because almost every program installation requires elevation in Windows.
> Macs have notoriously terrible hardware compatability
Hardware compatibility in what sense? As in, plug-and-play devices on a MacOS powered machine?
I'd argue the inverse; I often just plug in devices to my MacBook without having to install a single driver. Imagine my shock years ago when I plugged in a printer to my MacBook and I was able to immediately start printing without installing a single driver. Same with webcams, mice, etc.
Do you mean hardware compatibility in terms of build targets? I think here you might be correct, but even then you can compile for different operating systems from within MacOS... so again, I'm not entirely sure what you mean here.
I guess if you're talking about legacy devices where the hardware manufacturer hasn't bothered to create drivers for anything other than Windows, then your point might be valid, but how often does this happen...?
> Abstracting execution makes it harder to package software, harder to run it, and harder to debug it when something goes wrong
...more disingenuous statements. What do you mean by this? Under the hood, MacOS is Unix. Everything that runs on a MacOS machine is a process. You can attach to processes just as you would on a Windows machine. Similarly, if you have the debugging information for a binary you can inspect the running code as well.
MacOS is not a perfect operating system; for one, I do wish that it was better for game development. But I'm really struggling to understand your points here. Every single one is either not applicable or just straight up wrong.
macOS doesn't use the BSD VM.
> If anything, it's GPU is significantly weaker than most other laptops on the market.
You meant to say "stronger" at the task of not burning your lap.
It's just doable when you care enough to optimize the OS for responsiveness.
I'm actually not sure why it has the effect it does, entirely, because at least in terms of raw resource use the browser doesn't seem to be eating that much stuff, but the effect in practice has been well beyond what can be explained by perceptual placebo. (Maybe there's too many things competing to get the processor at the VBlank interval or something? Numerically, long before the CPUs are at 100% they act contended, even in my rather simple setup here.)
Or, to perhaps put it another way, Linux already has this sort of prioritization built in, and it works (though it's obviously not identical since we don't have split cores like that), but it seems underutilized. It's split into at least two parts, the CPU nice and the IO nice. CPU nice-ing can happen somewhat automatically with some heuristics on a process over time, but doing some management of ionice can help too, and in my experience, ionice is very effective. You can do something like a full-text index on the "idle" nice level and the rest of the system acts like it's hardly even happening, even on a spinning-rust hard drive (which is pretty impressive).
This, automatically turning on battery optimizations when an internal battery is detected, and automatically applying a 'small speaker' EQ to internal speakers are a few changes that would make Linux feel so much better on a laptop.
EDIT: I searched a bit and, although I haven't found instructions anywhere, I did find people saying that the low-latency kernel decreases throughput as a tradeoff to lower latency. Have you found this to be the case?
Also, another source says that the preempt kernel might be better for some workloads: https://itectec.com/ubuntu/ubuntu-choose-a-low-latency-kerne...
Can anyone comment on these?
Turns out that the cost of NUMA balancing (https://www.kernel.org/doc/Documentation/sysctl/kernel.txt) was outweighing any benefit we might get from it. The problem is that NUMA balancing is implemented by periodically unmapping pages, triggering a page fault if the page is accessed later. This lets the kernel move the page to the NUMA node that is accessing the memory. The page fault is not free, however; it has a measurable cost, and happens even for processes that have never migrated to another core. The lowlatency kernel turns off NUMA balancing by default.
Instead of switching to the lowlatency kernel, we set `kernel.numa_balancing = 0` in `sysctl.conf` and accepted the consequences.
In general terms, one either just installs the specifically recompiled package (-preempt, -lowlatency etc.), or recompiles their own. The related parameters can't be enabled, because they're chosen as kernel configuration, and compiled in.
I'd take such changes with a big grain of salt, because they're very much subject to perceptive bias.
I think most people just don't care or don't know that they care about responsiveness.
It seems like the same strategy would also make sense on Intel processors, although it probably requires at least 4 cores to make sense?
Intel agrees!
You also want them to be low-power even when saturated, otherwise you gain responsiveness from the "performance" cores but your "efficiency" cores aren't actually efficient.
It seems at least Linux's x86_energy_perf_policy tool lets you set multiplier ranges and some performance-vs-power values per-core, which means such a setup doesn't seem impossible on current Intel hardware.
As an example, a while back I ran a multi-threaded process on a shared work server with 24 cores. It used zero I/O and almost no RAM, but had to spend a couple of days with 24 threads to get the result. I ran it inside "chrt -i", which makes Linux only run it when there is absolutely nothing else it could do. I had someone email and complain about how I was hogging the server, because something like 90% of the CPU time was being spent on my process. That's because their processes spent 90% of their time waiting for disc/network. My process had zero impact on theirs, but it took some explaining.
The conclusions seem a bit off though
-Low QoS tasks are limited to certain cores. This doesn't necessitate efficiency cores (though it makes sense if you want power efficiency in a mobile configuration) and they could as easily be performance cores. The core facet is that the OS has core affinity for low priority tasks and quarantines them to a subset of cores. And it has properly configured low priority tasks as such.
-It also has nothing to do with ARM (as the original title surmised). It's all in the operating system and, again, core affinity. Windows can do this on Intel. macOS/iOS has heavily pushed priorities as meaningful and important, so now with the inclusion of efficiency cores they have established their ecosystem to be able to use it widely.
Apple actually pushes low priority tasks to the efficiency cores for power efficiency reasons (per the name), not user responsiveness. So that MBP runs for longer, cooler on a battery because things that can take longer run on more efficient, slower cores. They do the same with the efficiency cores on modern iOS devices.
The "feeling faster" is a side effect. I am talking about that side effect which isn't ARM specific.
And FWIW, Intel is adding efficiency/"little" cores to their upcoming architectures. Again for efficiency reasons.
It sounds like upcoming Intel chips will have the ability to run efficiency-mode on some cores, at which case OS-level code to do shuffling makes sense.
Windows solves this by placing basically everything onto the LITTLE cores, and user interactions serve as a priority bump up to the big cores. The key (which is difficult to implement in practice) is that they pass that priority along to every other called thread in the critical path so that every path necessary to respond to the user can be placed on the big cores. This means that if you are waiting for a response from the computer, every thread necessary to provide that response, including any normally low-priority background services, is able to run on the big cores.
I'd expect all the other OS's do a similar optimization for the big.LITTLE architectures. It seems to be the natural solution if you've worked in that space long enough.
[1] https://opensource.apple.com/source/xnu/xnu-4903.241.1/osfmk...
That's not bad!
However the M1 is only about 20-25% of the power consumption of the Ryzen 7 5600X.
That's crazy!
It'll be very interesting to see what the M3 or M4 look like (especially in X or other higher end variants) several years from now. We already have tastes of it from Marvell, Altera, Amazon and others, though it looks as if we will see ARM make huge headway in the data center (it already is) and even on desktops before 2025.
Granted I put my work laptop under more load (compiling, video chats at work) but it just feels like a normal laptop. I feel it take its time loading large programs (IDEs) and the fan clicks on when I'm doing big compiles.
I use my personal laptop for occasionally coding, playing games, and also video chats. But the M1 feels amazing. It's fast, snappy, with long battery life, and I've had it for several weeks and I never hear the fan. Even playing Magic Arena AND a Parallels VM runs silent. Arena alone turns my wife's Macbook Air into a jet turbine. It makes video chats so much nicer because it's dead silent.
I've run into occasional compatibility issues with the M1. sbt not working was the latest bummer.
So while my two laptops are technically a year apart, they feel like 5 years apart. The M1 laptop lives up to the hype and I'm glad to just have a laptop that's fast, quiet, has great battery, and is reliable.
Edit: More context, I had bought a maxed-out Macbook Air last year hoping it would be my "forever-laptop" but it was just not giving me the speed I wanted, and the noise was just too much. I couldn't play any game or do any real coding in the living room without disrupting my wife enjoying her own games or shows. I'm so glad I traded up to the M1
That is because there is no fan.
It's an M1 Macbook Pro, so there is a fan, but I've yet to hear it.
My previous laptop was an Intel Macbook Air (2019?).
Other than that it's been entirely silent and blazingly fast and I had no major issues with it.
Been a few years since I worked on this so I'm a bit fuzzy on the details but it's interesting stuff! Without this kind of tuning, modern flagship devices would run like absolute garbage and have terrible battery life.
Google Play likes to do shit in the background. A lot of it. All the time. Even if you have <s>deliberate RCE</s> auto-update disabled in settings, it will still update itself and Google services, silently, in the background, without any way to disable that. And while it's installing anything, let alone something as clumsy as Google services, your device grinds to a halt. It sometimes literally takes 10 seconds to respond to input, it's this bad. It's especially bad when you turn on a device that you haven't used in a long time. So it definitely wasn't scheduling background tasks on low-power cores, despite knowing which tasks are background and which are not (that's what ActivityManager is for). My understanding from looking at logcat was that this was caused by way too many apps having broadcast receivers that get triggered when something gets installed.
Now, in Android 11 (on Pixel 4a), this was partly alleviated by limiting how apps see other apps, so those broadcasts don't trigger much anything, and apps install freakishly fast. That, and maybe they've finally started separating tasks between cores like this. Or maybe they started doing that long ago but only in then-current kernel versions, and my previous phone was stuck with the one it shipped with despite system updates.
As for a particular strategy I presume Apple could, for example, use 6 high-performance cores instead of 4+4 hybrid. But then thermal management will be an issue. So the choice was about throughout/latency/energy efficiency. Using simple low-performance cores for tasks that can wait is very good strategy as one can put more of those cores as opposite to running the performance core under a low frequency.
Why is this? I understand that Intel chips don't use BIG/little but couldn't you assign all OS/background tasks to one core and leave the other cores for user tasks? Shouldn't the scheduler be able to do this?
macOS doing this is a side effect of them needing to do so, with the added benefit of it actually making things nicer all around. They could very well do this on intel/amd chips to the same effect.
The bummer is, we really don't know how good the silicon is, because nothing but macOS runs on it natively, so you can't really get an apples to apples comparison.
With slow light power cores, that's less of an issue.
No idea how anyone works in an office with one of those things without irritating the hell out of their coworkers.
There were even people saying they deliberately sabotaged their intel thermals to make the up and coming M1 look better.
I'm not familiar with Android APP, maybe there are similar APIs in Android?
My point is, the snappiness is a result of decent system-wide adherence to a priority API more than the heterogeneous chip design. The low power consumption of course relies on both.
But you then get a huge tradeoff: If you reserve over 50% of the computational power for "asap" stuff, then all the "whenever" stuff will take twice as long as usual (assuming the tasks to be limited by compute power). On a 8C Intel that means that, for a lot of stuff, you're now essentially running on a 4C Intel. Given the price of these chips that MIGHT be a difficult sell, even for Apple. On the M1 they don't seem to care.
We could call it ::SetThreadPriority()
Intel does have a chip for you though, they have a version with mixed UlV cores and big cores.
It would be a bit of a hassle to figure out which processes would all need to be changed, and whether these changes persist after rebooting and there's probably a lot more that I didn't think off though..
https://www.windowscentral.com/assign-specific-processor-cor...
My 2018 USB-C MBP honestly feels like trash at times, just feels like every part of it is choking to hold the thing up.
The Time Machine backup pictured above ran ridiculously slowly, taking over 15 minutes to back up less than 1 GB of files. Had I not been watching it in Activity Monitor, I would have been completely unaware of its poor performance. Because Macs with Intel processors can’t segregate their tasks onto different cores in the same way, when macOS starts to choke on something it affects user processes too.
this is really smart, and makes one wonder why intel didnt have the insight to do something like that already.... is this the result of thier "missing the boat" on smartphone processors (a.k.a giving up xscale)??Corollary question: could one make a kernel-specific core and would there be benefit to it? Handle all the I/O, interrupts and such, all with a dedicated, cloistered root-level cache.
If the goal is low latency then staying on a single core is almost always better. To effectively use multiple cores requires explicit synchronization such as io_uring which uses a lock-free ring buffer to transfer opcodes and results using shared buffers visible from userspace and the kernel. io_uring has an option to dedicate a kernel thread to servicing a particular ring buffer, and this can also be limited/expanded to a set of cores. I have zero experience with io_uring in practice and so I don't know what a good tradeoff is between servicing a ring buffer from multiple or single cores. The entries are lightweight and so cache coherency probably isn't too expensive and so for a high CPU workload that also needs high throughout allowing other cores to service IO probably makes sense.
I think newer x86_64 chips also allow assigning interrupts from specific hardware to specific cores to effectively run kernel drivers mostly on a subset of cores, or to spread it to all cores under heavy I/O.
https://developer.apple.com/library/archive/documentation/Pe...
It used to be that I would notice rogue background tasks eating up my processor because I could feel the UI acting slow and the fans would kick on.
Now the only indication I get is my battery depleting more rapidly than expected. The UI is completely responsive still, and there are no fans. But I’m never checking my battery level anymore, so I don’t catch these processes for much longer.
https://en.wikipedia.org/wiki/PA-RISC
That's an interesting comparison because it's another case where a big player created their own chips for a competitive advantage.
There are surely some lessons in there for Apple silicon, but I don't there there's any particular reason to think it will follow the same path as HP PA. PA started in the '80s. It was aimed at high-end workstations and servers. My impression was that the value proposition was narrow: it was sold to customers who needed speed and had a pretty big budget. Until checking now, I was unaware, but it seems to have carved out a niche and was able to live there for a while, so it wasn't exactly a failure.
Time will tell how apple silicon fares. Today the M1 is a fast, power-efficient, inexpensive chip, with a reasonable (and rapidly improving) software compatibility story, and is available in some nice form factors... so basically a towering home run. But we'll just have to see how Apple sustains it over the long run. Frankly, they can drop a few balls and still have a strong story, but it's hard to project out 5, 10, 20 years and understand where this might end up.
(Personally, I'm almost surely going to get an Apple silicon Mac, as soon as they release the one for non-low-end machines. The risk looks small, the benefits look big.)
This means they do not have to be dependent on Mac sales to support M1 development - they can amortize it over the largest phone manufacturer in the world.
And every time they do a successful architecture switch they make it easier to do it again in the future. If the IntAMD386 becomes the best chip of 2030 they can just start using it.
I often have the feeling on my i9 machine it's almost idle but the specific core of my UI is working on something else which I don't particularly care for at the moment I'm interactive with the computer.
I get that not everyone is a "take matters into their own hands" kinda person, but it's worthwhile to at least look into it and see if it's something worth pursuing for you.
Anyone know why this is? seems fairly arbitrary...why not have 1,2,3,4,5?
An AMD/Intel using same soldered RAM next to CPU and same process node would give Apple a run for its money.
Still, the optimizations on OS side are interesting here.
CPU and GPU and AI share unified memory which means zero copy transfers between them.
AMD, Intel, ARM, & Qualcomm have all been shipping unified memory for 5+ years. I'd assume all the A* SoCs have been unified memory for that matter too unless Apple made the weirdest of cost cuts.
Moreover literally none of the benchmarks out there include anything at all that involves copying/moving data between the CPU, GPU, and AI units. They are almost always strictly-CPU benchmarks (which the M1 does great in), or strictly-GPU benchmarks (where the M1 is good for integrated but that's about it)
> An AMD/Intel using same soldered RAM next to CPU and same process node would give Apple a run for its money.
AMD's memory latency is already better than the M1's. Apple's soldered RAM isn't a performance choice:
"In terms of memory latency, we’re seeing a (rather expected) reduction compared to the A14, measuring 96ns at 128MB full random test depth, compared to 102ns on the A14." source: https://www.anandtech.com/show/16252/mac-mini-apple-m1-teste...
"In the DRAM region, we’re measuring 78.8ns on the 5950X versus 86.0ns on the 3950X." https://www.anandtech.com/show/16214/amd-zen-3-ryzen-deep-di...
Careful what you are comparing, in your examples the other CPU is also faster.
3950x is a desktop CPU and is faster then M1 -> https://gadgetversus.com/processor/apple-m1-vs-amd-ryzen-9-3...
5950x is even faster -> https://gadgetversus.com/processor/apple-m1-vs-amd-ryzen-9-5...
Lower latencies are likely due to higher clock.
For equivalent laptop specific CPU, you will get a speedup from on-package RAM vs user replaceable RAM placed further away, even desktops would benefit but it would not be a welcome change there.
That's not really how dram latency works. In basically all CPUs the memory controller runs at a different clock than the CPU cores do, typically at the same clock as the DRAM itself but not always.
If you meant the dram was running faster on the AMD system then also no. The M1 is using 4266mhz modules while the AMD system was running 3200mhz ram
> For equivalent laptop specific CPU, you will get a speedup from on-package RAM vs user replaceable RAM placed further away, even desktops would benefit but it would not be a welcome change there.
Huge citation needed. There's currently no real world product that matches that claim nor a theoretical one as the physical trace length is minimal latency difference and far from the major factor.
It's more accurate to say that M1 made the jump from the A14 chips already in iPhones, as it's an enhanced variation of that design.
Lots of Intel damage control these days. AMD is kicking their butt from one side and ARM from the other.
8GB M1 Mini, so maybe that makes the situation worse.
So do Slack and Spotify...
That's an 8 year old CPU! If you're having issues, the problem lies elsewhere not with Electron.
I don't have any measurements, and I'm not really going to take any. It's just always felt slow. So I'll admit I should've perhaps been less prescriptive.
It's been known M1 is super fast launching apps compared to Intel macbooks - https://www.youtube.com/watch?v=vKRDlkyILNY
There's tons of video comparisons on YouTube.
Perhaps you have in mind some specific problematic app?
Or you include Intel apps that are translated to ARM on the fly on their first launch?
I just described my bad experience with M1 with Big Sur. Not sure who is to blame.
Although I imagine given how slow the response was on the butterfly keyboard issue it might be a few years out.
There isn't one in M1.
Not an overheating issue, no, but the M1 still benefits from a fan, and obviously whatever Apple does in a Mac Pro class machine will also have a fan. They aren't going to keep at ~25W in a full size desktop, that'd be nonsense.
The M1 is designed for this. Wide cores that do a lot per clock, but can't be clocked high is perfect, because Apple would rather not clock high. Intel and AMD are still fighting the MHz wars, and can't release a core design that doesn't approach 5 GHz, even if the lower power chips won't.
What I can't stand are all the people saying the M1 is capable of replacing high end x86 computers. It's like owning a Tempo and mistakenly thinking your highly specialized transmission means you can walk up and challenge a 6.7L Dodge Challenger to a race. It's completely ridiculous and demonstrates a stunning lack of self awareness.
But the M1 is also literally, by the numbers, the fastest machine in the world for some workloads. For us, my M1 Macbook Air compiles our large TypeScript Angular app in half the time of a current i9 with 64Gb ram on node v15. These are real-world, tested, and validated results. Nothing comes close..because the chip is designed to run JS as fast as physically possible. It's insanely fast at some workloads. Not all of them of course, but it's no Tempo. I would say it's more like a Tesla, faster than it should be considering the HP (a model S is much faster than cars with way more HP) but faster due to 'How' it makes the power rather than how much.
That said, I replaced a $2000 iMac 4K with a base model Macbook Air and it was a huge upgrade for my day to day work. It really is perfectly fine to replace some workstations.
Essentially the same here, I just happened to replace my $2-3k Win10/Ubuntu desktop with my 8GB(!) M1 Air and it truly has been a huge upgrade for my day to day work. It feels like I'm working in the future.
That said - I primarily use this Air to ssh to all of my servers that do my heavy lifting for me. But that doesn't stop me from driving this thing hard - dozens of open tabs, multiple heavyweight apps left open (e.g. Office apps), multiple instances of VS Code, Slack, Teams, all running - zero slowdown. Zero fan sound.
It's black magic good.
I think you have other things going on that's slowing down the compilation on the i9.
Maybe if run linux and install a fast SSD your i9 will be faster than your m1. Then by your logic you should ditch the m1.
Whether or not this is possible completely depends on their unique workload. The M1 could probaly obviate any x86 chip for the average programmers. But that wouldn't be the case for gamer workloads, for example.
A high end gaming PC will compile typescript, play games with high quality graphics, and run your 100 browser tabs. Let's be honest, an Intel Core 2 Duo would obviate most programmer workloads just fine.
By your logic the Apple M1 is comparable to a high-end x86 chips, but you go on in the same breath with the disclaimer that it is "only certain workloads."
That is my point. A high end PC doesn't care about your workload. It is unobjectively fast. Period. You are the Tempo driver claiming that you would have beaten that Challenger if the temperature of the drag strip were just a little bit hotter....
But that Challenger doesn't care about operating conditions or workload or whatever. It will win today, tomorrow, and the next day while you're claiming that the Tempo is "comparable in the snow on Tuesdays."