Apple's Cyclone Microarchitecture Detailed
anandtech.com
anandtech.com
There are many reasons in favour, including keeping negotiations with Intel over price interesting, but probably the main thing preventing them running with it would be an inability to manufacture the chips fast enough, which is a major concern in mobile land. This is why lots of Android devices exist in variations based on different chips as hedges by the device manufacturer on long term availability.
Also, Apple probably have little problem getting preferential treatment from manufacturers.
Yep. IIRC Apple ships something like 10 million laptops/year. They ship 3 times that in phones… every quarter
Also remember Apple has already successfully made two or three architecture transitions already: m68k to PPC to i386 to x86_64.
It would be the third party applications that would be a hurdle. Without a transparent layer like the deprecated Rosetta, they would not be able to run existing applications and would have to explain that a new MacBook Air can't run all of the old MacBook Air applications.
Maybe if they come out with an ARM powered laptop, they will call it something other than MacBook Air to avoid that confusion.
Don't forget that Apple has spent years getting everyone on Xcode, and disciplining them to avoid the parts that would make a recompile & test straight-forward.
They also have a "good enough" office suite (for many people), as well as the iWork apps.
Such a device could potentially also run iPad apps without much change.
OS X and its predecessors have already been ported to run on a bunch of architectures: x86-64, i386, PPC64, PPC32, PA-RISC, SPARC, and Motorola 68k. All, or nearly all, of the CPU-dependent parts of the OS are the parts that are shared between OS X and iOS. Making OS X run on ARM would be trivial for Apple, and I'd be shocked if they didn't have it up and running already as a hedge.
If they go down this road, Apple will not put constraints like that on their Mac developers: I'd expect that all the most commonly used frameworks would continue to be available, and porting most applications would amount to a recompile against a new Xcode.
With a port to ARM, though, Apple can simply say that the developers haven't ported their app to the new architecture yet, and that more apps will become available as developers catch up. Meanwhile, all those apps that currently require some kind of lower-level access could be banned.
On the other hand, I think that would be a huge boon for non-power users; it would be nearly impossible for malware to get onto the computer, and even annoyances like Adobe's and Microsoft's auto-updaters would finally get funneled though the App Store, which would prevent programs like that from constantly occupying memory/CPU/network.
For this to work, though, I think Apple would be wise to include some kind of developer mode feature (perhaps even forking over the $99/year fee) that would allow unsigned or potentially dangerous apps to run.
All of this (admittedly wild; sorry, coffee is kicking in) speculation definitely has me excited for the future of Apple, though! I haven't felt like that since Steve was around.
Wacky. Your vision of shutting out unsigned apps has me looking for the exits. I've been an Apple user since the 80s and I'm pretty sure this would finally make me jump ship. Apple's current Big Brother trajectory frightens me to no end.
The notion that this CPU is really about giving ample power to the iOS line makes more sense to me. There will eventually be an "iPad Pro," whether or not it gets stuck with that moniker, and it won't have to worry about backward compatibility with anything other than previous iOS devices.
Especially because they were known to have OS X running on x86 as far back as 10.0, years before they actually switched to Intel officially.
This seems very likely to me. They presumably have a fair bit of the work done with iOS, which shares a kernel codebase and so on, and really, it'd be worth the effort just to have it to show to Intel executives during negotiation.
So, it's mainly a matter of porting the higher level stuff -- mainly the Cocoa libs, which are slightly different on iOS vs desktop OS X.
This would've been the perfect phrase last year.
I do however, think that they will eventually include both arm and x86 processors in the Macbook Air. That way, backwards compatibility is preserved and low power apps can run on the arm. In the current Macbook Pros they have dynamic switching of GPUs, there's no reason they couldn't use the arm as a coprocessor or even run the full OS.
Here a few technical points:
* LLVM - You can compile once for both both architecture and then the JIT will take over compiling for the specific architecture (Take a look at llvm for the OpenGL pipe line in OSX)
Full screen app- When an app is full screen (if it's a low power app) then the x86 could sleep and the arm could switch to running the app.
App nap - If all x86 apps are a sleep, switch over the running the arm processor exclusively
*Saving State - It's possible to save the state of apps in OSX, a similar mechanism could be used to seamlessly migrated between processors.
This is pure speculation, but it is feasible. There would be many technical challenges that Apple would have to solve but the a capable. The advantage Apple has is that they have absolute control over both platforms.
At one point I had a machine which allowed movement of data between apps running in Windows 3.1 on a DX4 and Risc OS on a StrongARM. I'm wondering how hard it would be with OSX (a far stricter OS than Win3.1 or Risc OS) to allow a sort of process level separation, where the OS and any ARM compatible processes are running on an ARM, but you keep an x86 around and spin it up for certain processes.
| Data | ARM insructions |
You could certainly remove the ARM instructions and replace them with x86 instructions. However, the ARM instructions will have hard-coded certain offsets in the data buffer, like where to look for global variables. You would have to be sure that the x86 instructions had exactly the same offsets. For another issue, if the data buffer contains any function pointers, then the x86 and ARM functions had better start at exactly the same offsets. And if there are any alignment requirements that differ between x86 and ARM (I don't know if there are), then the data had better be aligned to the less permissive standard on both chips.None of these problems are impossible to solve. They could be solved easily by adding a layer of indirection, at the cost of some speed, and then Apple could go back and do the real difficult-but-fast implementation later.
However, why would it? When its ARM cores are essentially desktop-class, there's no need to have an x86 chip other than compatibility with legacy code. Looking at Apple's history, it seems pretty clear that it likes to have full control of its own destiny, and designing its own chips is a logical part of that, so having its own architecture could be considered a strategic move too.
So given the difficulty of implementing it well, and assuming that Apple eventually wants to have exclusively Apple-designed ARM chips in all of its products, if I were in their shoes, I wouldn't bother to make switching work. I might have a product with both kinds of chips, but I would just have the x86 chip turn on for x86 apps, and off when there were no x86 apps running, and know that eventually those apps would go away. (And because I'm Apple, I have no problem pushing vendors to switch to ARM faster than they want to, so this won't be a long transition.)
However, an even cooler move would be to make LLVM IR the official binary representation of OS X, and compile it as part of the install step of a new program. That gives Apple several neat capabilities:
1) They can optimize code for the specific microarchitecture of your computer. Maybe not a huge deal, but nice
2) They can iterate on their microarchitecture without having to care about the ISA, because the ISA is an implementation detail. This is the technically correct thing that everyone should have done years ago (yes, I'm annoyed).
3) They can keep more secrets about their chips. It's obnoxious, but Apple would probably care about that.
So, there's my transition plan for Apple to move to its own chips. It probably has many holes, but the biggest one is still the question of what Apple gains from this. Intel still has the best fabs, and as long as that's true, there will be some advantage in sticking with them. Whether the advantage is big enough, I don't know. (And when it ends in a few years, then who knows?)
I've wondered the same thing. In this respect the IR is analogous to Java bytecode, or C# CLI. In this context these share conceptual similarities, allowing for multiple languages to target the same runtime.
This would possibly open up iOS to being able to be more easily targeted by languages-that-aren't-Objective-C. As long as it compiles down to LLVM IR then this "binary" becomes language agnostic. (Actually, for all I know things like RubyMotion do this today. I haven't delved into it to find out.)
So a user installing Firefox or Chrome or some other complex application would need to wait for tens of minutes before they can use their application? It's more likely they'll just re-use the existing dual arch architecture but instead of PPC/x86 it'll be x86/ARM…
It's worth revisiting http://lists.cs.uiuc.edu/pipermail/llvmdev/2011-October/0437... ("LLVM IR is a compiler IR"), from a core LLVM developer, explaining why LLVM IR is unsuitable for this task.
Given such a tool, which existed in 1992, it seems simple enough to do the recompile once on the first launch and cache it. Executable code is a vanishingly small bit of the disk use of an OS X machine.
Going forward, Apple has a long experience with fat binaries for architecture changes. 68k→PPC, PPC→IA32, IA32→x86-64. I don't think x86-64→ARM8 is anything more than a small bump in the road.
As far as shipping LLVM and letting the machines do the last step, that should make software developers uncomfortable. Recall that one of the reasons OpenBSD needs so much money² for their build farm is because they keep a lot of architectures going because bugs show up in the different backends. I know I want to have tested the exact stream of opcodes my customer is going to get.
␄
¹ I think there was also a MIPS to Alpha tool for people coming from that side.
² In the sense that some people think $20k/yr for electricity is a lot.
Why? This is how Windows Phone 8 works (MDIL) and Android will as well (ART).
On WP 8 case, MDIL are ARM/x86 binaries just with symbolic names left in the executable. The symbolic names are resolved into memory addresses at installation time, by a simplified on device linker.
Android's ART, already made default on the public development tree, compiles dex to machine code on device installation.
Going forward, Apple has a long experience with fat binaries for architecture
changes. 68k→PPC, PPC→IA32, IA32→x86-64. I don't think x86-64→ARM8 is anything
more than a small bump in the road.
Using the lipo[0] tool provided as part of the Apple Developer tools, it's pretty easy for any developer to create a x86/ARM fat binary. Many iOS developers have used this technique to create libraries that work on both the iOS simulator as well as a iOS device.There was a paper at ASPLOS 2012 where they did something like this, but for ARM+MIPS [1]. Each program would have identical ARM and MIPS code (which took some effort), with identical data layout.
The IR isn't architecture portable right now. IE, you can't use it as a live interpreter language, because the code it produces make assumptions on the target architecture before final binary translation.
It would be fantastic if Apple would fix LLVM so the IR was portable, it would be amazing for general purpose software if you could ship LLVM IR and have your end users compile it or have web services do it for target devices on demand.
I do however, think that they will eventually include both arm and x86
processors in the Macbook Air. That way, backwards compatibility is
preserved and low power apps can run on the arm. In the current Macbook
Pros they have dynamic switching of GPUs, there's no reason they couldn't
use the arm as a coprocessor or even run the full OS.
Battery life of the Macbook Air is already outstanding. I think I want something more than just better battery life for all that complexity, but eventually it will all be just "taken care of" by the toolchain, so why not?I think the low hanging fruit would be running iOS apps native on the Macbook Air.
It's not you're emulating an entire operating system: the operating system (and many libraries) are native, but the application code is emulated. It's faster than you think, and Apple has already done it twice: once in 1994, and once in 2005 (exercise for the reader: try extrapolating).
Apple's applications would be 100% native long before the ARM version shipped. Some intensive tasks—text rendering, audio/video/image encoding and decoding, HTML layout, JavaScript—these would also be native on third-party apps, since you just have to write the right glue into the emulator. This would be a lot easier than the 2005 switch from PowerPC to x86, which involved emulating a system with 4x as many GPRs and the opposite endian: ARM has 2x the GPRs as x86-64.
Sure, a bunch of apps will see reduced performance. Some will break. But remember: Apple has only been on x86 for ten years. We had the same problems during the PowerPC->x86 transition: you had to wait about two years to get a version of Photoshop that ran on x86 + OS X.
I'm willing to bet that Apple has been testing OS X on ARM for years now.
Well in PowerPC->x86 transition x86 was the faster chip. The emulation cost was discounted by some amount. If you go from x86->ARM, ARM is the slower chip, so there's never going to be _improvement_ in performance compared to x86. I don't see why you're equating the two.
The whole XCode toolchain already supports x86_64 and arm64. It's trivial to build "fat" binaries with clang that run on both platforms.
The switch from PPC to x86 was less painful than many thought. I don't see any particular reason why the "fat binary" approach (which is technically gross, but quite practical) would not work for switching from x86 to ARM.
Developers will hate it, but the whole app store farce shows what Apple thinks of developers.
https://github.com/llvm-mirror/llvm/blob/7b837d8c75f78fe55c9...
http://lists.cs.uiuc.edu/pipermail/llvmdev/2014-March/thread...
https://github.com/llvm-mirror/llvm/commit/7b837d8c75f78fe55...
This is truly amazing. In comparison, Cortex-A15 can issue 3 instructions per cycle, for example.
That said, I'm going to be watching for the Denver chip, it will be an interesting counter point to the A7.
Essentially, Qualcomm's market.
Cortex A-15 _does_ run at nearly double the clock rate (though with a substantially higher energy consumption on the rare occasion it actually scales up all the way). Apple is historically very conservative with its mobile device clock speeds, and I'd expect that Cyclone has a fair bit of headroom here.
http://en.wikipedia.org/wiki/ARM_Cortex-A15#Systems_on_a_chi...
Configure an ARM platform with 8 GB of DDR3 DRAM and a PCIe-based SSD, and you may well blow the power budget versus the current Intel platform design.
Intel has turned its attention to total platform power: moving the VRMs on-die, and also identifying third-party motherboard components which draw silly amounts of power.
I think that Intel is on track here. Apple will of course push their design talent in order to deliver the best mobile devices they possibly can. But this does not necessarily mean a move away from Intel for a laptop or desktop.
For example see the second chart here: http://preshing.com/20120208/a-look-back-at-single-threaded-...
The memory bottleneck is the real problem, but once you start getting systems conceptually related to nVidia's Tegra K1 (or even the PS4 system architecture) where the CPU is relegated to a sort of system co-ordinator Intel are going to have a major headache and find themselves in that single threaded performance niche that kept the SPARC and POWERs of the world occupied prior to their decline.
(Yes, I know marketing says "ad spend" but it's still wrong)
nouninformal noun: spend; plural noun: spends 1. an amount of money paid for a particular purpose or over a particular period of time. "the average spend at the cafe is about $10 a head"
I'm sorry.
Sooner or later mobile processors are going to hit the power wall. I haven't looked into this lately but one thing that could provide a fixed and permanent benefit to mobile processors is the more modern instruction set.
Especially the first P4s, those were awful. A Tualatin P3 @ 1.2/1.3GHz would be much faster than a P4 @ 1.4 or even 1.6
And those Northwood cores reached 3.4GHz, compared to 3.8GHz of the infamous Prescott with it's 31 stages. That is fairly meager clockspeed increase, and based on that I'd argue that the extreme pipeline depth was not the major contributing factor in pushing the clockspeed of P4 higher.
Not so long ago I remember doing some calculations on the SPEC results, normalising them to a single thread and clock frequency to determine the per-cycle efficiency of various CPUs, and x86 came out at least a factor of 2-3x ahead of the rest.
From a marketing point of view, too, it'd be a hard sell in Android-land. Apple has been very careful to steer clear of spec-oriented marketing, but can you imagine the less-sophisticated enthusiast market's response to, say, the Galaxy S6 using a dual-core 1.3GHz chip instead of a quad-core 2GHz?
Plus, from a market standpoint, Cyclone was developed to meet Apple's own needs. To sell these chips would introduce a "demand" variable into the equation that I think would stifle development.
At any rate, this is a really interesting topic because I think that Apple's oft-critized isolation actually worked much to its own benefit here.
Didn't AMD, and Intel more or less say the same thing about 3 wide? No real benefit from going wider. Is that because of the differences in micro architecture or is it more about getting a little more performance without having to ramp the clock speed? What makes 6 wide good for ARM but not x86 or x64?
Apple have made several aquisitions in the last few years to give them the technology they need to make these sorts of developments. Even if some of the tech had been stolen, it would require a lot of work to put it into practice.
I know I'm feeding the troll, but I couldn't help it; this is just too ridiculous.
Unfortunately the ARM architecture, even with those optimizations is probably slower "clock by clock" compared with x86/PPC
But yes, I think this is something Apple is probably testing (Desktop Mac OSX on ARM)