Apple’s Mac Chip Switch Is Double Trouble for Intel
bloomberg.com
bloomberg.com
I have three feelings:
1. Apple will just enable the existing chips “today” and developers can ship ARM code to a sizable audience, with instant battery benefits.
2. Cheaper Macs will start dropping x64 next year and continue running only ARM code, maybe emulating x64 like Rosetta.
3. Pros will keep having both processors for the foreseeable future, similarly to how they have “dual graphics”, except that the OS will always run on ARM.
We should keep in mind that Apple just dropped a swath of 32-bit software which means they aren’t afraid to do it again.
But maybe the situation has evolve ? Expired patents or a deal between Apple and Intel ?
Supporting a different ISA on the same OS/Runtime is relatively “easy” in 2020 and doesn’t need such a sophisticated strategy - for 99 percent of developers this is going to be a one click recompile to different build target in their IDE. Push new binaries to App Store and done.
By “easy” I mean relative to how ISA transitions have been managed in past. Apple are so big now that runtime mechanisms for x86 code, such as Rosetta during the PowerPC to x86 migration or hardware emulation on the 68k etc, just aren’t necessary this time around IMO. People will recompile. For sure we will lose some software that is too old to be recompiled by long lost authors, but I don’t think Apple cares. They already just killed 32bit app support with similar consequences.
Remember today that your “arm” iOS apps work just fine in your x86-based iOS simulator in Xcode on an intel Mac, so access to actual ARM hardware isn’t a significant barrier at all to development for most.
1. If x64 wasn't already dominant and had the same market share as ARM, would you choose x64? Why?
2. With x64's current dominance, would you buy a ARM system as your primary one, knowing that you'll have significant difficulties developing for the majority of current market share?
Patents have a 20 year lifespan, most of the x86_32 ones have expired, but x86_64 is just barely still covered (the first publication was in 2000, and first commercial product was on AMD Opteron/Athlon64 in 2003). So it's probably just a matter of time before the userland emulation stuff is further extended, although we won't get anything newer than SSE2 for the same reason.
I think this is a very small issue.
The reason why it's hard to develop for ARM today is that reasonably powerful (read: non-phone) ARM processors are not very common. The barrier for developing on ARM is that you have to explicitly purchase another ~$500 laptop in order to test on ARM. Also, libraries generally support x86 but not necessarily ARM. With small desktop market share, the expense is often not worth it.
On the other hand, x86 processors are ubiquitous, and if you're a developer, you can almost certainly find a old laptop or desktop with an x86 processor laying around to test your software on (especially if you're in the demographic to buy relatively expensive MacBooks). Also, since most people use x86, it makes sense to invest in that market.
TL;DR: ARM -> x86 barrier to entry is much lower because x86 processors are so common, x86 opportunity is much more lucrative, and most libraries already support x86.
Here is one example of what you get for 50 dollars:
https://www.hardkernel.com/shop/odroid-c4/
This is a quadcore 2GHz 64bit board that can run Android / Ubuntu 20 / Wayland / WebGL / Linux Kernel 5.4 / Vulkan.
I would pick the one which performed the best on my workload.
Assuming everything was equal; performance, performance per watt, software support, market share, price to performance: I would pick ARM because I think the competitive landscape is more diverse. x86 is effectively an AMD-Intel duopoly, especially with the cross-licensing agreements.
It's a lot easier to start working with ARM IP, which means there should be more competitors, which should translate to better future performance improvements. Even today, we see many more companies competing in this space (Ampere, Marvell, Qualcomm, Apple to a degree), not to mention startups looking to develop their own ARM chips.
[1] https://techcrunch.com/2020/04/29/arm-is-offering-early-stag...
I develop on a Windows laptop and deploy to Linux all of the time. I have some Python programs that have native dependencies. I still develop on Windows, push and the CI/CD pipeline runs on Linux and packages a Windows build.
But even with my first job out of college back in the 90s, I was writing C code that I developed on Windows and was cross compiled for DEC VAX and Stratus VOS mainframes.
Should be “packages a Linux build”
“ I was writing C code that I developed on Windows and was cross compiled for DEC VAX and Stratus VOS mainframes.”
Of course cross compiles is not the correct terminology. We built the same code on the target machines.
Mainframes do it, as means to integrate C and C++ into their language environments.
Back in the early mobile OS wars, there was a company selling J2ME like stack, but using C and C++ instead.
Then there is the LLVM bitcode used by Apple on iOS and watchOS (which happens to be more platform neutral than regular LLVM bitcode).
Oh and WebAssembly and MSIL as well.
LLVM proprietary Apple own internal fork, as used in iOS and watchOS, is another matter.
As a matter of fact, there is a WWDC talk about how it allowed the seamless migration of 32 to 64 on watchOS.
https://mobile.twitter.com/clattner_llvm/status/104696072464...
Apple does whatever they feel like with their proprietary fork, to the point that Apple's clang also gets its own column on cppreference.
Just like Sony and Nintendo haven't contributed anything back to LLVM that would disclose any capability from their consoles.
I tweeted @atp and let them know that the transcript was returning a 404. They have since fixed it.
https://atp.fm/205-chris-lattner-interview-transcript
John Siracusa: The same thing I would assume for architecture changes, especially if there was an endian difference, because endianness is visible from the C world, so you can’t target different endianness?
Chris Lattner: Yep. It’s not something that magically solves all portability problems, but it is very useful for specific problems that Apple’s faced in the past.
Only peripherally related to your point, but I develop with Python on Windows and found that deploying to Linux, while in possible, can be a hassle in practice.
For instance, I use turbodbc on Windows to access SQL Server databases via ODBC. This works fine.
However, when I try to port the same code to ARM Linux (on Raspberry Pi), I learn that ODBC has dependency on the native ODBC driver which doesn't exist on the ARM platform. So I have to jump through all kinds of loops to compile FreeTDS and unixODBC, which aren't trivial. Not only that, I had to cross-compile on an x64-Linux VM to the ARM architecture because the Raspberry Pi itself didn't have enough disk space for gcc and such.
I think cross-architecture compilation isn't the issue -- it's cross-platform deployment, especially when there's no parity in dependency availability.
1) You're almost certainly decoding into u-ops even if you chose RISC because you'll have uarch features like a seperate pipe line for the AGU and the load store queues, atomics that have to wait to go out all the way to L2 for fairly arbitrary lengths of time, etc. You can see this in cores as simple as BOOM, and it's opinion of the RISC-V community that macro op fusion of prescribed clauses is the way to go.
2) These decoders are a drop in the bucket when compared to OoO circuitry and power budget.
3) The complex addressing modes and memory RMW operands are effectively a way to address physical registers while consuming no architectural registers, and very few bits of I$. Yes, x86 is ancient and isn't as optimal as it could be from a huffman encoding perspective (hlt is a single byte opcode!), but it's pretty damn good overall. Better than AArch64 at code density and therefore I$ pressure. As an aside, I'm sorta curious what a CISC-V would look like, and if it would set a new bar.
I've been toying with the idea of literally having decoding as decompression, where there's a special instruction to change the dictionary. I guess this'd be tantamount to implementing the decoder as an FPGA, but I'm hoping there's some reasonable version where a fairly non-dense "base encoding" becomes a pretty optimal bit stream.
That being said, my experiments were hardly conclusive and I'd absolutely love to be proven wrong.
More specifically a core where the heart is being sequenced by Tomasulo's algorithm, and probably a large bypass network linking the functional units together.
Try doing that with ARM and tell me under what limited circumstances would you buy ARM laptop/desktop/workstation.
CPUs are at a stage where AMD and Intel are doing great in non battery constrained setups and with battery constraints you will only get better battery life if you don't make the CPUs sweat. In server/workstation design I would expect AMD and Intel to beat the crap out of ARM anything just because how much optimized they are in that space owing to massive use.
So except for frothing at mouth idiots that keep droning about this there isn't much value for the normal customer in buying ARM anything and lose on all fronts - software availability, build flexibility, compatible hardware availability and even performance for many cases. (Spare me teh Geekbench please.)
It does make sense for Apple to market their own CPUs to people who are into that type of thing - it's more control, more money and least dependence for Apple - all round win for them if they can pull it off. But it's not going to be easy unless they have something that is hugely better and overcomes at least some of the disadvantages.
I think this might be one of the motivations for Apple to switch to it's own CPUs instead of just switching to a different x86-64 CPU like AMD's Ryzen. Apple would be able to lock MacOS to the hardware like iOS and the iPhone.
Download apps only from Mac App Store, make it easy for people 100% in Apple ecosystem to do what they need to do (XCode, browsing, store apps, multi tasking) and they're good to go. They already charge premium for the hardware - not having to pay Intel will fatten the margins, with app store revenue and maybe banning 3rd party browsing engines they can get exactly what they have always wanted and so long as they provide people with some macOS features - terminal access (even if controlled) and multitasking etc. - people in the Apple ecosystem aren't going to care.
Of course it maybe gradual and it maybe not so restrictive - they may allow 3rd party app installations and make it harder like they have already done in a way.
Anandtech:
"This year, the A13 has essentially matched best that AMD and Intel have to offer "
https://www.anandtech.com/show/14892/the-apple-iphone-11-pro...
All at, as far as I can tell, a small fraction of the power consumption / thermal load.
So imagine a featherweight 12" MacBook with roughly the CPU performance of a much larger/heavier Intel laptop and the same or better battery life.
"This year, the A13 has essentially matched best that AMD and Intel have to offer – in SPECint2006 at least."
So yeah your comment is business as usual for Apple fan - taking comments and benchmarks out of context, comparing apples to oranges and just being generally dense to support some narrative. I.E. Nothing closer to the reality of having ARM CPUs beat AMD or Intel if I wanted to build a server or workstation or even a big powerful laptop with not optimized for email and browsing battery life.
SPEC doesn’t agree with your characterization of their benchmark.
“The SPEC CPU® 2006 benchmark is SPEC's industry-standardized, CPU-intensive benchmark suite, stressing a system's processor, memory subsystem and compiler.”
Neither does Wikipedia:
“SPECint is a computer benchmark specification for CPU integer processing power. It is maintained by the Standard Performance Evaluation Corporation (SPEC). SPECint is the integer performance testing component of the SPEC test suite”
https://en.wikipedia.org/wiki/SPECint
Nothing about energy efficiency, and notice the definite article “the” in front of “integer performance testing component”.
So SPECint tests general compute tasks, such as compiling, XML processing and running Perl programs.
Of course you can combine the perf measurements of SPECint with power consumption measurements to arrive at an efficiency measure, if you so choose.
The other benchmark was SPECfp, which focuses on scientific computing tasks such fluid dynamics, quantum chemistry etc.
So tasks you are unlikely to perform on...your iPhone. The A13 still does well, but not as well as on integer, presumably because Apple just didn’t focus on putting idle resources on their iPhone chip.
Increasing FP performance when you have great integer performance is straightforward, afaik: just add more FPU resources.
But leaving that aside, it is also a single threaded performance measure which is nice but not quite a big deal for desktop / workstation / server workloads. It's not very convincing that they will have a 12 core part for workstation that I can use (leaving aside the fact that I won't be able to run what I want on it) - that will perform much better than AMD/Intel in real world workloads including virtualization etc. So I still maintain that narrowly focused performance benchmarks Geekbench or SpecInt mean very little in factoring OP's question on whether you will replace your primary system with ARM - the other CPUs are also limited to what Apple's CPUs can run - not Java stuff for example.
ARM as a whole had lot of time to compete with others in desktop/workstation/server market and there is little evidence they got anywhere big.
It is simply the business model changes. The older I get the more I am convinced it isn't the technical that changes or determines the outcome, it is the value proposition.
Not to mention no one "choose" x86. They chose Intel, or Intel CPU that still currently offer top performance under 10 - 16 Core.
No offense, but the libraries Intel is putting out there (IPP, MKL) for efficient parallel computing are really outstanding. It has been a fee years for me in the end-consumer high-performance market, but AMD chips would easily run 20% slower if you optimized the code w/ the Intel Libraries and OpenCL would Not get anywhere near.
A similar thing happened w/ Nvidia and Apple many more years back; E.g., Apple dropping Nvidia; fair amount of rumors went around that Apple did that so Adobe would be less competitive if they had to rewrite their rendering engine (which they just ported to CUDA) and Apple could sell their video editing solutions with an edge (they have been hand tuned for a while of course...).
Apple dropping Nvidia has been such a “non-customer-focused” bullshit decision. Not because the AMD hardware is bad - but OpenCL has simply been nowhere near CUDA. Same with Intel IPPs and OpenCL. But maybe that changed.
The differences in performance, ease of use etc. have been mind-blowing back in the days for CUDA and IPP. Hardware alone isn’t gonna cut it. The machines are built to run software after all...
Maybe that has changed or maybe nobody needs any parallelised compute intensive applications on Mac (research anyone, image and video editing anyone?).
The fact aside that mayor software companies like Adobe et al. will probably have to staff entire departments for rewrites...
There's a full XNU based OS on there called "BridgeOS" that's in the iOS/tvOS/watchOS family.
However, there were PPC upgrade cards for 68K Macs but you had to reboot to switch from 68K to PPC.
https://www.engadget.com/2020/02/07/apple-may-testing-amd-pr...
Even more out there idea (with literally zero proof, it's just a good idea, IMO): Apple buys Centaur, and gets an x86 licence.
* Apple gets to have custom power efficient cores augmented by all the fabless firms they've acquired over the years.
* Mac stays x86, so Intel wins a minor victory when the alternative was a major customer switching to ARM.
* Intel wins a bigger victory because an x86 licence is effectively removed from products being on the open market. The great equalizer that is the end of Moore's law makes that licence sitting out there a long term existential risk for Intel.
* Apple doesn't have a costly transition with the tail end of Moore's law meaning they don't have the same perf gains expected from the other transitions.
* Apple also puts their hands on Centaur's newer inference accelerator IP.
* Centaur's parent keeping them on life support gets a payout.
Everyone wins except AMD (which is another win for Intel).
So it would make no sense to trade Intel's poor schedule and performance for AMD's uncertain future performance.
And aren't all x86 licensee prevented from keeping that license if they are bought?
Tim Cook was talking about context nearly a decade old.
> So it would make no sense to trade Intel's poor schedule and performance for AMD's uncertain future performance.
x86 is just going to become more of a commodity as time goes on. And currently, Zen 2 is hands down the best perf/watt combo currently.
> And aren't all x86 licensee prevented from keeping that license if they are bought?
That was a the rumor, but centaur has already been bought and kept it's license, so at a minimum there there's some fine print to that clause. And my experience with B2B is that clauses like that are ultimately a product of the circumstances from when they're written. If circumstances change, those clauses can change. The most indelible ink is the most likely to have new semantics later.
No x86, no virtualbox. No virtualbox and I'd rather just get (with much sighing, complaining, and general ill will) a windows laptop. I mean, the terminal is becoming slightly more useable in windows, right? And sometimes you just need to run a windows VM, so you need x86 virtualization.
There are downstream effects to losing the "nerd" base and I suspect this move is a bad idea, but apple has a track record of pulling rabbits out of hats and knowing what really matters. Maybe this customer segment just doesn't matter.
I don't feel the same using WSL. It's not a seamless experience. Also, I've never found a terminal emulator I enjoy using on Windows. Maybe it's still way too early, but even the new Windows Terminal left me less than impressed.
Also, is the Mac Pro going to switch to ARM? I'm not aware of an ARM chip that can compete with the super highend Xeons in the Pro. Having laptops run ARM and Pros run x86_64 doesn't seem like the best idea (also sounds like a lot of work on Apple's part).
Of course, maybe this switch is going to create a high-end ARM space, allowing ARM to make inroads into HEDT and the server market.
A lot seems unclear at the moment, but one thing is clear (to me atleast): there's going to be a huge fight over the next 10 years, x86_64 vs ARM. No one can possibly know who will win, but it's exciting to the say the least. I think we've all been a little tired of the x86_64 monoculture since the end of PPC.
A lot of open source code works well on ARM. But will we start to see some newly-discovered-but-latent arch-specific bugs, compiler bugs, undefined-behavior bugs-but-worked-on-x86_64 bugs? Yes, sure.
The cool thing is that Win/ARM and Linux on arm are still very much the same OS as their x86_64 ports. Presumably macOS is/will be the same way.
ARM gets less love precisely because they're not as popular for developer native workstations. But I wouldn't be surprised if that changes over the next decade.
They will be able to hit desktop -> mobile in one shot.
It took developers a long time to have stuff fully ported, but it wasn't a huge deal - if you switch to a chip that's twice as fast and take a 50% hit on the emulation layer, you haven't lost much.
I'm very curious to see stuff that isn't first- or second-party already-natively-recompiled stuff works. There's a lot of quality-of-life Mac native apps that I like, that aren't backed by a bunch of dev resources.
The PPC Macs couldn’t emulate a 68K floating point unit at all.
Almost nobody (except imgix and a few others) run macos on a prod server, yet many devs run macos. For example: when they run stuff via docker they run it via a VM (whether they know it or not).
Any dev (again, except imgix and a few others) that actually cares about server performance is already not running their benchmarks/perftests/tests on a mac, so that should not make a difference.
A lot of code for demanding applications is often enhanced with SSE/AVX and JIT recompilation techniques. Those are inherently unportable, and I'm not sure how many developers will be willing or able to port that code over to Neon and AArch64, especially for a small 2-4% of the market.
Even if they do, it's quite a cognitive burden to have master and maintain two separate SIMD and recompilation implementations for the same applications.
This will probably further drive high-end gaming away from Macs, on the heels of OpenGL deprecation and the Mac-only Metal API. Combined with Cocoa and Swift, I imagine we'll end up seeing less and less applications that run natively on both Macs and Windows/Linux after the move.
Agree. I guess people will end up using Stadia/GFN/PS Now, etc.
There is sse2neon https://github.com/jratcliff63367/sse2neon. For intrinsics it supports, you only need to add a header. There is also simde https://github.com/nemequ/simde. It is a larger project and may be more complete.
iOS developers?
Couple of benchmarks show intel clearly winning in the multi-core world but it looks like single core performance is a bit more of a tossup.
https://gadgetversus.com/processor/apple-vs-intel-core-i9-99...
One thesis of the article is that people will start to deploy on ARM (e.g. Gravitron) so that it will be the same architecture.
The vast majority of well written C and C++ code will compile and just work on ARM64 as long as it doesn't depend on things that are undefined in the C/C++ spec but X64 lets you get away with. The big bugaboos are unaligned memory access, assembly or X64 intrinsics, and reckless casting. Vector code will need porting (assuming there's not already a NEON version) but that's generally only found in graphics, audio, machine learning, and cryptography applications. Most apps don't have any of that. (Auto-vectorization is irrelevant as the compiler does that.)
Code in higher level or newer languages is generally even less worrisome. Rust, Go, Java, C#, and any dynamic language will just work.
I do expect that these ARM64 chips are going to have lower per-core single-threaded performance but higher parallel performance than X64 due to more cores. That means that some applications may need refactoring or partial redesigns to be more parallel to take full advantage of the chip. But that's something that needs to happen anyway since all architectures are going many-core due to the end of big easy single core performance gains. It's been a long time since huge gains in single threaded performance were a thing.
As long as they don't screw it up by e.g. nerfing the OS I'm looking forward to better battery life and better overall performance due to many cores and lack of X64 instruction decode bottlenecks.
I'm curious about how they'll do it though. I predict instruction translation (X64->ARM64), but also the return of fat binaries. It's also possible that apps distributed through the App Store will be delivered only in your host architecture automatically. I think they were making some noise about that a while ago and that may be prep for this.
iOS apps—which run on ARM chips inside iPhones—have all been developed on Intel-based Macs.
I guess we'll see soon enough. I doubt they're as naive as to believe that they can pull another massive arch switch without Jobs' reality distortion field, and without their current arch severely lagging (like it was in PPC->x86 transition).
And additionally with the latest iPad becoming much closer to ‘Mac’ in terms of the browsing experience (eg with a touchpad) it makes total sense for Apple to start making some of their apps truly cross platform. Does the Spotify desktop app really need to be substantially different to the iPad app?
I'm a little surprised to see Bloomberg basically cribbing MacRumors.
EDIT: Instantly downvoted? LOL