Apple just demonstrated that changing CPU architecture is not that big of a deal; so the value of x86 compatibility isn't what it used to be. You can just emulate it and it's fine. Even for games apparently. So Intel, backing an architecture that is already starting to compete with arm that is free makes a lot of sense.
AMD has the same challenge. And despite Nvidia failing to buy ARM, it's pretty clear that their long term strategy is not going to be letting other companies supply CPUs but to provide a complete solution.
The amount of raw effort that Apple put into making that transition _appear_ “no big deal” will be hard to adequately appreciate! It _noticeably_ affected software quality at least two MacOS versions prior (drop of 32-bit support and forcing all API accesses to go via their frameworks) and will have cost them hundreds of millions if not billions in engineering effort and unknowable amounts of lost sales in the mean time.
Yes it paid off. Obviously. But emulating their move, I’m not sure if there is even a single company able to do that.
https://www.zdnet.com/article/intel-we-have-arm-license-no-p...
That's if you have monopoly like control over your entire eco-system...
If the Apple fandom wiki is correct, the 68k emulator for PPC was included in all PPC releases, but Rosetta was included in 10.4 and 10.5, optional in 10.6 and unsupported in 10.7; it's scope was more limited than the 68k emulator as well. I expect Rosetta 2 will have a similar limited lifetime.
"Everybody" knows that Apple has an ARM architectural license, but AFAIU the terms haven't been disclosed. Presumably they got a sweeter deal than other ARM architectural licensees when they got rid of their ownership in ARM ages ago, but, "don't have to pay a thing" and "can do whatever they want" sounds a bit too sweet to be true?
> So, developing risc v in the background without committing to it, yet, makes a lot of sense.
TBH, I think Intel's interest in RISC-V is more about hurting ARM in the embedded market than about planning to sunset x86.
Oh yes, absolutely. (I was going to mention that earlier, but the edit timer had expired.)
But yes, splitting fab service into a separate business unit that is seriously open for 3rd parties seem to be a major strategic shift since Gelsinger took over the helm. And it probably makes sense, as TSMC et al have demonstrated the merchant fab model can work for the top end designs as well and the entire rest of the industry is moving towards that.
So in a way, unless Intel wants to be the odd man out with their own idiosyncratic workflows this is a route they must go down on.
Not quite.. Apple also had to modify ARM on the hardware side to support x64's stronger memory ordering. But RISC-V has options for both I think.
For an attempt at doing it without modifying the hardware see Microsoft's slow emulation attempt on their ARM version that was rejected by consumers and increased battery consumption much more.
Software designed for the normal RISC-V memory model will work perfectly on a TSO machine, if perhaps a bit more slowly. Programs that depend on TSO semantics may be buggy on the standard RISC-V memory model (or on ARM too).
RISC-V also has a FENCE.TSO instruction that can be inserted as needed into software running on normal RISC-V. If you're going to use it a lot then you'd be better off implementing an actual TSO mode (not least because of code size). FENCE.TSO even works on (standards compliant) hardware that doesn't know about it, because unknown fences are supposed to be executed as FENCE RW,RW (the strongest fence), at some loss in efficiency.
Alibaba T-Head unfortunately didn't read this part of the spec when they designed the C906 and C910 cores, which give illegal instruction trap if they encounter an unknown fence such as FENCE.TSO. OpenSBI now does trap-and-emulate if necessary, but of course at another loss of efficiency. This bug affects the Allwinner D1 and Alibaba ICE SoC.
I don't know whether this errata in the C910 has been fixed in the TH1520 SoC in the Roma laptop (and other unannounced, cheaper, SBCs).
License is the least of it.
Apple released their first custom chip in 2013. And spent several years before that designing it.
So, it took Apple up to 15 years to get to where they are now with M1.
Even if AMD and Intel start now, and are twice as fast at developing new CPUs, they will need 6-7 years to reach parity with 2022 Apple in... 2028
A new competitive processor in an area where all you have is "some experience"... Well, I wouldn't hold my breath.
If (and that's a big if) AMD and Intel started looking into ARM seriously after Apple unveiled M1, I wouldn't expect any ARM processor out of them earlier than 2025-2026.
Funnily enough I'd expect Amazon to perform better in this space (Graviton has been in deployment since 2019, three years ago, and is now in its third iteration).
So AMD hasn't had much in the ARM department in the past 5-7 years.
My guess is that they kept an ARM Zen frontend working, internally. And they probably have a RISC-V frontend now, alongside many other projects. They are large enough to do so.
When they perceive they can launch a successful product, they do so. Otherwise, we never know of these efforts.
Sure, the battery life and not overheating is great, but you are limited to Arm based solutions, when installing linux VMs for example. Or gaming. Everyone seems to have forgotten how awesome it was to be able to do everything on one machine.
It is a compromise I live with, not something I prefer.
And at least half of Apple’s benefit comes from process technology, where the gap will be closing very soon, if for no other reason than the fact that after 2nm there is not much more room to grow with silicon, so even laggards will have time to catch up on process and yields.
The would has been running on x86 for so long it’s a waste to just decide to rework so many chunks of it.
Just need to find a way to justify it to myself after spending on the MBP xD
Apple's performance-per-watt advantage is boosted by process improvements, but nothing I've read attributes anything close to half of that advantage to process. (References would be welcome.)
"Chips manufactured with this [3nm] process deliver 30% improvements in power consumption and 15% better performance in comparison with 5nm chips." https://www.computerworld.com/article/3609778/apple-on-track...
Where they fall down is on perf/watt, not raw performance. However there’s a lot of things that go into that difference, not just ISA, so I’m not sure if anyone has really decided if x86 is fundamentally less efficient, or just currently less efficient due to current design choices and constraints.
Raptor Lake is a pretty big jump in perf/watt. In just one generation they are claiming similar perf to Alder Lake at 250W but just 65W on Raptor Lake. AMD did a similar huge jump in efficiency last year with Ryzen 6000 for mobile
No. AMD and Intel are the oligopolistic providers for an unbelievably vast software ecosystem that practically rules all computing outside embedded (and some legacy mainframes here and there), are they going to throw away that market position just because the cleaner encoding of ARM or RISC-V would save an estimated low single-digit % of decoding power [1]?
[1]: https://www.usenix.org/system/files/conference/cooldc16/cool...
This problem does not apply to RISC-V, where with the C extension you get either 32bit or 2x 16bit. The added complexity is negligible, to the point where if a chip has any cache or rom in it, using C becomes a net benefit in area and power.
ARMv8 AArch64 made a critical mistake in adopting a fixed 32bit opcode size. A mistake we can see in practice when looking at the L1 cache size that Apple M1 needed to compensate for poor code density.
L1 is never free. It is always very costly: Its size dictates area the cache takes, the latency of this cache, the clocks the cache itself can achieve (which in turns caps the speed of the CPU), and how much power the cache draws.
As you mentioned decoder width: There's Ascalon[0], a RISC-V microarchitecture that's 8-decode (like M1), and 10-issue, by Jim Keller's team at Tenstorrent. It isn't in the market yet, but is bound to be among the first RISC-V chips targeting very high performance.
Note that, at that size (8-decode implies lots of execution units, a relatively large design), the negligible overhead of C extension is invisible. There's only gains to be had.
C extension decode overhead would only apply in the comically impractical scenario of a core that has neither L1 Cache nor any ROM in the chip. Such a specialized chip would simply not implement C. Otherwise, it is a net win.
Micro-op caching enters the room.
I mean, seriously, micro-op caching has been extensively used for over 20 years, including on ARM64 designs although Apple M1 doesn't have it.
In any case, if Intel needs to diversify, open source (RISC-V) does seem better than proprietary mortal enemy (ARM).