>That whole "optimized for macOS" thing is a myth. Heck, we don't even have CPU deep idle support yet and people are reporting 7-10h of battery runtime on Linux. With software rendering running a composited desktop. No GPU.
https://twitter.com/marcan42/status/1498923099169132545
Also, from a couple of days ago, they got basic suspend to work.
>WiFi S3 sleep works, which means s2idle suspend works!
The M1 has a fairly good GPU, so there's hope that the battery life and overall experience will improve in the future. As of now though, I'd reckon there's dozens of x86 Linux laptops that can outlast the M1 on Linux. Pitting a recent Ryzen laptop versus an Asahi Macbook isn't even a fair fight.
https://dougallj.wordpress.com/2022/04/01/converting-integer...
https://dougallj.wordpress.com/2022/05/22/faster-crc32-on-th...
https://lemire.me/blog/2020/12/13/arm-macbook-vs-intel-macbo... (I later optimised the slower benchmark in that post: https://github.com/simdjson/simdjson/pull/1708 )
Obviously the GPU will be better, but at one point I compared the M1 CPU to other ARM GPUs (in laptops at that time) and found it had both better memory bandwidth and compute throughput, which is quite funny.
GPU input programs can be expensive to switch, because they're expected to change relatively rarely. The vast majority of computations are pure or mostly-pure and are expected to be parallelized as part of the semantics. Memory layouts are generally constrained to make tasks extremely local, with a lot less unpredictable memory access than a CPU needs to deal with (almost no pointer chasing for instance, very little stack access, most access to large arrays by explicit stride). Where there is unpredictable access, the expectation is that there is a ton of batched work of the same job type, so it's okay if memory access is slow since the latency can be hidden by just switching between instances of the job really quickly (much faster than switching between OS threads, which can be totally different programs). Branching is expected to be rare and not required to run efficiently, loops generally assumed to terminate, almost no dynamic allocation, programs are expected to use lower precision operations most of the time, etc. etc.
Being able to assume all these things about the target program allows for a quite different hardware design that's highly optimized for running GPU workloads. The vast majority of GPU silicon is devoted to super wide vector instructions, with large numbers of registers and hardware threads to ensure that they can stay constantly fed. Very little is spent on things like speculation, instruction decoding, branch prediction, massively out of order execution, and all the other goodies we've come to expect from CPUs to make our predominantly single threaded programs faster.
i.e., the reason that GPUs end up being huge power drains isn't because they're energy inefficient (in most cases, anyway)--it's because they can often achieve really high utilization for their target workloads, something that's extremely difficult to achieve on CPUs.
This part here 100x. It's worth noting that the SIMD performance of the M1's GPU at 3w is probably better than the M1's CPU running at 15w. It's simply because the GPU is accelerated for that workload, and a neccessary component of a functioning computer (even on x86).
The particularly damning aspect here is that ARM is truly awful at GPU calculations. x86 is too, but most CPUs ship with hardware extensions that offer redundant hardware acceleration for the CPU. At least x86 can sorta hardware-accelerate a software-rendered desktop. ARM has to emulate GPU instructions using NEON, which yields truly pitiful results. The GPU is a critical piece of the M1 SOC, at least for full-resolution desktop usage.
[0] https://asahilinux.org/2021/08/progress-report-august-2021/#...
And yet, after first release use, they have reported that this is the smoothest they've seen a Linux desktop ever run. That is, smoother even when compared to intel-Linux on hw GPU acceleration.
Even extremely fast CPUs suck really bad at pushing pixels compared to even the weakest GPUs. It is very much usable though!
Currently llvmpipe is able to use up to 8 cores at a time but not more, and does use SIMD instructions when available from my understanding. There is another software rendering system in MESA that allegedly uses AVX instructions but I have had a better experience with llvmpipe personally.
I have already been reading the state of running x86 vms on macos on an m1, it is slow but my friends claim usable. Add it on an early linux impl on native m1, then x86 vm.
Why am I torturing myself? I want to get back to having a great fanless system like I used to have on a google pixel laptop, where they had underclocked x86 but running without a fan was great.
https://developer.apple.com/documentation/virtualization/run...
Since Rosetta for Linux was coerced to run on non-Apple ARM Linux systems almost as soon as it shipped in developer preview builds of the upcoming macOS, it would not be surprising if Asahi ends up making use of it outside of a VM context (though that may not comply with the license).
And this seems to be able to get it running: https://github.com/diddledani/macOS-Linux-VM-with-Rosetta
I haven't tested this because I'm at work, but I'll verify it when I get home!