I mean, seriously, all that tradition and experience and you have a phone company make circles around you on your own field.
I mean, seriously, all that tradition and experience and you have a phone company make circles around you on your own field.
Apple is in a unique position of being able to force a new architecture on its customers, without losing them. They have done it twice. They even aren't exactly compatible with normal ARM, due to a special agreement with ARM Holdings.
Intel had I960, quite cool and successful, and could not capitalize on it in the long term for low-power devices. Intel bought rights to ARM, and could not capitalize on it in the long term either, even though ARM was well-suited for battery-powered devices, and sold it!
Intel used to be the king of data centres, using an architecture from 1980s, extended and pimped up to the brim — but still beholden by backwards compatibility. And a king it still is. But this pillar seems to shake more and more.
I don't think this is a result of poor engineering. It was, to my mind, a set of business bets, which worked well, until they didn't any more.
I read the rest after typing this.
Intel was...it was king.
To be honest - Apple actually got some pretty damn good performance out of the PowerPC chips and architecture - my Quad-Core G5 tower with 16GB RAM is still used for finalization of my music projects, due to its insanely smooth performance - and tbh coming from a very experienced user of modern Macs it still kicks ass.
It is true that Apple implemented a bunch of custom optional features (some of which are, arguably, in violation of architectural expectations), and they definitely have some kind of deal with ARM to be able to do this, but from a developer perspective they are all optional and can be ignored. I don't think Apple exposes any of them directly to iOS/macOS developers. They only use them internally in their own software and libraries (some are for Rosetta, some are used in Accelerate.framework, some are used to implement MAP_JIT and pthread_jit_write_protect_np, some are only used by the kernel).
They very clearly have a special, Apple-only relationship with ARM.
FWIW, the compresssion instructions are used by the kernel, and I don't even know if they work from userspace. I've only ever tried them in EL2.
https://developer.arm.com/documentation/ddi0595/2021-03/AArc...
Apple's, in marketing speak, brand permission has given them extraordinary latitude over the past 15 to 20 years. They've been able to get off with making abrupt transitions and other relatively wrenching choices that tech pubs and doubtless forums like this wailed about but which their customers were mostly fine with. Things that Microsoft and WinTel laptops, for example, couldn't with respect to ports, limited options, etc. couldn't.
The main challenge is seamlessly migrating users to the new platform and Apple did a great job at it using Rosetta.
Both Linux and especially Windows struggle at this because they lack something as well integrated as Rosetta and require all applications to be recompiled to a new architecture.
As a result Microsoft has to worry about what the OEMs want and how they will use the product. In contrast Apple only cares what the end user wants, what their experience is, what features they get and how they work.
An OEM cares about whether they're making ARM laptops or Intel laptops. They care about and want input into the implementation details. An end user doesn't care if Photoshop is running on an ARM chip or an Intel chip, they care about how well it runs and what the battery life is. They (generally speaking) don't care about the implementation details.
Far less than a lot of people probably think though. AMD has an excellent x86 core, faster single threaded and throughput than the M1, on a generation older process technology, and quite possibly a smaller design and development budget than Apple, although not so power efficient.
Microsoft tried that with the Surface X and failed
When Apple rolls out a product, there's some transitional overlap, but you can see them getting ready to burn their viking ships in that period. Microsoft's efforts always have lacked that kind of commit factor. IMO.
Lack of key software from Microsoft doomed them to a niche.
People say this without thinking. There is no real evidence at all it is true.
Something like x86 support on an IA64 chip costs extra transistors. But there is no real fundamental reason why it should make anything slower.
This is even more so for AVX512 instructions, which aren't backwards compatible in anyway.
So - exactly - how would dropping backwards compatibility speed up AVX512 division?
But that's one switch, implemented in hardware in the decode pipeline.
It makes implementation more complicated, but no reason it has to be slower.
Now I wonder why x64 can't re-encode the instructions: Put a flag somewhere that enables the new encoding, but keep the semantics. This would make costs for the switch low. There will be some trouble, e.g you can't use full 64bit values. But mostly it seems manageable.
This is incorrect.
Intel Skylake has 5 parallel decoders (I think M1 has 8): https://en.wikichip.org/wiki/intel/microarchitectures/skylak...
AMD Zen has 4: https://en.wikichip.org/wiki/amd/microarchitectures/zen#Deco...
Which is impressive, if you think about it. But it is also complicated machinery for a part that's basically free when insns are fixed width and you wire the buffer straight to the instruction decoders. Expanding the pre decoder to 32 bytes would take a lot of hardware, while fixed width just means a few more wires.
https://stackoverflow.com/questions/23788236/get-size-of-ass...
The same technique could be extended to cover all of them and and it's not so difficult to implement this in verilog.
As long as this state machine runs at the same throughput as the icache bandwidth then it is not the bottleneck. It shouldn't be too difficult to achieve that.
But it is definitely extra complexity, and requires space and power.
Then again, both Intel and AMD make it work, so there must be a way, if you're willing to pay the hardware cost. Now I think about it, the same linear to logarithmic trick for adders can be done here: Put a state machine before every possible byte, and throw away any result where the previous predecoder said skip
This also demonstrates where it really hurts is when you want to do something low cost, and very low power, with a small die. And that's where ARM and RISCV shine. The same ISA (and therefore toolchain, in theory), can do everything from the tiniest microcontroller to the huge server. This is not the case for x86.
Strap a lot of cash and volume behind that after they get bought, and the results speak for themselves. There is no shame in it. Intel has been stumbling at the moment, but others were stumbling in the decade from Core 2 to Skylake.
PA Semi aren't a phone company, they just work for one ...
"These guys aren't just going to walk in and ..."
(famous last words, LXXXV)
and
"x86 has been a continuously supported backwards compatible architecture for 35 years, and enabled the existence of most computers for most of those 35 years"
are basically saying the same thing. Depending on which way you look at it, Intel can feel bad about it, or feel good about it.