ARM64 and You
mikeash.com
mikeash.com
For those interested, there's an another side. A64 drops all the features of the ISA (inline variable shifts, conditional execution, variable-width instructions) that are hard to implement in a fast, high-power CPU. If a cpu is to not have any 32-bit ARM compatibility, there's no reason one couldn't make a 4GHz 4-wide superscalar one based on A64.
And in any case: there's another reasonably well-known company out there with an even cruftier ISA making 4+GHz 6-wide superscalar cores with backwards 32 (and 16!) bit compatibility modes. Instruction decode is the sort of thing programmers understand, so we tend to get hung up on it when talking about CPUs. It's really not a meaningful design limitation to a CPU core implemented with hundreds of millions of transistors.
Apple's track record proves that they are more than happy to drop anything more than a couple of years old if it means general forward progression.
My real bugbear is that to some extent they've created the modern day IE6 by not updating Safari on the old iPad.
Also, if you look at some legacy iOS apps on the Appstore (designed pre 6.x), they are pretty much 100% broken on 6.x+
Not from Apple, maybe. I know there is at least one in the works for the server market.
> And in any case: there's another reasonably well-known company out there with an even cruftier ISA making 4+GHz 6-wide superscalar cores with backwards 32 (and 16!) bit compatibility modes. Instruction decode is the sort of thing programmers understand, so we tend to get hung up on it when talking about CPUs.
You absolutely can implement a fast core out of almost any ISA (just look at IBM System z!) if you have sufficient resources. The point is that A64 is amenable to being made into a fast CPU without Intel-scale design effort.
As others mentioned, though, I don't think Apple will need to do that since they've got a closed ecosystem. They can just declare one day that they won't accept App Store submissions that don't include a 64-bit version and very quickly move the entire universe of iOS software. There will be some no-longer maintained apps that aren't updated but they can just say "Sorry, this app doesn't support iOS 8"
There are also binary-translation techniques like what they did with Rosetta in the PowerPC->Intel transition. This wouldn't need to be done on the phone; they could just do the translation once on the app-store side.
When i was learning x86 assembly, discovering CMOV was fantastic, drastically simplified all my code (versus cmp+je+hundreds of extraneous labels.. notwithstanding macros). The fact ARM could do that for almost all instructions was one of the main reasons in my mind why ARM was considered a "cleaner" architecture than x86.
EDIT: crisis averted, a comment further down ( https://news.ycombinator.com/item?id=6458457 ) clarifies the change.
Some, yes (but it is retained in 32 bit mode).
The reason it was dropped is that it's hard to implement. It adds an implicit data dependency between an instruction and previous results. In a pipeline, the results of previous instructions aren't available immediately, there is usually a delay of many cycles before a "previous" instruction actually finishes and retires. So if your "current" instruction needs the result of the one before it, it's going to have to wait, stalling the pipeline. Adding extra dependencies such as instructions that depend on the state of the flags register is not desirable. IIRC it also makes register renaming hard when one implicit register (the flags) is heavily used.
It appears that they've been using tagged pointers on the desktop since 10.7, which I never realized: http://objectivistc.tumblr.com/post/7872364181/tagged-pointe...
Interesting that Apple now has the coordination and control to make it a non nightmare thing now.
Apple has done some work to alleviate this extra memory pressure at the kernel level. grep for WKdm in the xnu sources if you're interested.
[1] http://groups.csail.mit.edu/cag/raw/documents/R4400_Uman_boo...
I've only just started looking at AArch64 but I agree that it feels a lot more like MIPS though. I think that's a good thing.
Probably because so many projects use Thumb (the default for iOS projects in XCode, for example) which doesn't include most instructions for conditional execution. From what I can tell, it also sounds like compilers weren't making very effective use of those instructions anyway.
Also, these were originally meant to compensate for a lack of branch prediction, which as I understand it, has changed drastically in recent years.
but http://www.arm.com/files/downloads/ARMv8_Architecture.pdf says:
31 general purpose registers accessible at all times * Improved performance and energy
* General purpose registers are 64-bits wide
* No banking of general purpose registers
* Stack pointer is not a general purpose register
* PC is not a general purpose register
* Additional dedicated zero register available for most instructions
Which one is it?
By the way, the ARMv8 resources are quite interesting overall and a bit more in-depth than the article. http://www.arm.com/products/processors/armv8-architecture.ph...
It sounds like you're partially confusing this with r9, which is a completely different story - r9 was globally reserved by the system until iOS 3.0, then allowed. This equivalent in arm64 is x18, which again is globally reserved by the system.
More generally, a frame pointer of some sort is necessary to deal with variable-sized stack allocations, and often useful for performance analysis and debugging tools, so many compilers set them up by default even when they aren’t strictly necessary.
From the 31 GPRs, x30 is the link register and x29 is the frame pointer. x30 can be used for other purposes within a routine, but x29 must always hold a valid frame record in the iOS ABI. Additionally, iOS reserves x18 (“The platform register”) for all use. So there are really 28 GPRs, or 29 if you include x30/lr, which is something of a hybrid.
I've had a few quibbles about where performance gains would be, and all too often I was told that the performance increases would be solely realized in the larger memory addressing space. That just didn't seem right to me.
I really like the use of the otherwise unused space in the 64-bit pointers.
A64 doesn't eliminate conditional execution completely. It just pares it down to the basics: branch (obviously), add/sub, select, compare (for flattening conditionals like `a && b && c`).
Another thing removed from A32 was the optional shift on operand 2-- which was taking up 7/32 bits for most instructions.
This has a few more that were missed: http://nominolo.blogspot.com/2012/07/arms-new-64-bit-instruc...
You can use this pattern to implement compare-and-set, but you don't need compare-and-set to use that pattern directly.
Edit: I wasn't sure how to encode this into the steps in the article, so it's a bit vague on that part. Suggestions welcome.
I expect we'll see ARMv8 architectures in the next round of flagship phones. Apple's a little ahead of the curve, but it won't be long till competitors catch up.
In the context of Apple, it's interesting to think about how they're going to take this next. ARM process and architecture improvements are likely to lead to chips with high-enough performance to be used in mainstream desktop applications – Is it possible we're going to see something like an ARM/x86 dual-processor Macbook platform that allows ARM's low power consumption supplement Intel's performance?
ARM64 Apple chips are play for iPads. Current iPads are lagging on performance when we compare it to something like a Baytrail Intel tablet or a Haswell equipped surface tablet. There is going to be convergence point for Intel where a tablet with Haswell level performance with a Fanless chasis and 500$ price. Apple need to converge there to compete.
Explicit APIs can have explicit error codes. Memory accesses don't have much opportunity to report errors, so have to resort to awful signals and such (that nobody handles properly).
On the subject of memory mapping and magnetic disks, one amusing bit of history is that GNU's Hurd kernel originally implemented filesystems by memory mapping the entire hard drive and working from there. This worked fine at first, but started to cause major trouble when HDs grew beyond 4GB and Hurd was still running on 32-bit CPUs. I believe they ended up redoing it all without memory mapping so they could grow beyond that limit.
2) They took the money they saved and devoted it elsewhere (perhaps the Sapphire home button? :)
When it comes down to the bill of materials, every cent really does count when it scales across several million units sold.
I personally think Apple strategically held off the RAM upgrade 'til next year's iPhone, so that they could have a "killer feature" to lean on if they don't manage to figure out a more novel^Winnovative one in time.
I could easily see a situation where a bunch of people were sitting around saying, "y'know, we don't absolutely positively NEED a 64-bit processor, but it doesn't cost much more and there might be some performance gains and if nothing else it might be a good bullet point for marketing. On the other hand, doubling the RAM will double what we pay for RAM, it won't help performance, and it will eat into the power budget, and it wouldn't make a good bullet point for marketing. The only downside is that denim_chicken will think we're being nefarious and will tell the world about it on HN."
I have no idea if that's what happened, but Occam's Razor suggests that it's more likely than the combination of pure evil and pure incompetence that you've been postulating here.
The power consumption angle is a complete and utter non-issue. It generally comes up as an apologetic canard to justify Apple's choice here, but in the holistic sense the power difference between 1GB and 2GB is negligible. Note that the ARMv8 processor, however, is a serious power pig.
However the limited memory should be an issue for people. I personally seriously considered a 5S to replace my GS3, largely for the fantastic camera, but the window of credible life for the 5S is simply too short -- 1GB just isn't enough, and seems especially deficient compared to such a fantastic processor.
This is AnandTech's iPhone 5s battery review: http://www.anandtech.com/show/7335/the-iphone-5s-review/9
The 5s outperforms iPhone 5 in four out of five tests. Regression was only seen in one scenario. The thing is, I don't remember coming across any smartphone review, iPhone or otherwise, that single out the CPU as the source of power consumption change. It's always a combination of different factors.
Now it's totally possible that the new CPU is a power pig as you claimed. Unless you can provide proofs that validate it, though, I have to agree with others that you're making it up.
The same review that notes the increased power consumption of the CPU? That one?
The iPhone wifi, display, and surrounding platform is identical to the iPhone 5. The LTE/3G chipset is improved (not surprising as it's a considerable power consumer, which was why Apple held out on LTE for a while). On the wifi test, where all else is the same as before, the iPhone 5S saw a 10% longevity decline despite a 10% larger battery.
Quite humorous seeing so many so desperately defensive about this, when the original (and completely unsubstantiated) claim was that going to 2GB would be see a marked increase in power consumption. We know from these very results that you provided that the device did see a 20% or more decrease in longevity, mAh to mAh, despite the fact that the CPU is generally a small consumer of power (for the whole device to consume 20% more power, the CPU had to have increased significantly more). Power "pig" is relative, and obviously it's a ridiculously low power processor by any normal metric, but compared to the one they replaced it with...yeah.
What's your point? It's a beefier SOC, and possibly more power hungry. Nobody disputes that.
You've tried to frame rebukes of your posts as "desperately defensive". For me I take issue with what you said below:
> The power consumption angle is a complete and utter non-issue. It generally comes up as an apologetic canard to justify Apple's choice here, but in the >holistic sense the power difference between 1GB and 2GB is negligible. Note that the ARMv8 processor, however, is a serious power pig
You seem to be absolutely sure, with numbers to back it up. You dismiss the OP as "completely unsubstantiated", but ironically you can't prove your points either. Should I take you seriously?
It's a significant fraction of total power draw when asleep. You also have to account for PMU efficiency being pretty low when running at low current, so minor changes are an even larger difference. There's plenty of publicly available figures you can find to research this - find the "IDD6" self-refresh current in an lpDDR datasheet. Scale to the size of DDR you want, compare with battery rating adjusted for voltage.
So a device with 2GB should have a significantly smaller standby time than one with 1GB, right, given that, by your claims, it's a significant fraction. Only that isn't true at all, and straight comparisons between, for instance, the GS3 international (1GB) and the GS3 Snapdragon (2GB) shows absolutely no reduction in two-week+ standby time. Of course all else isn't the same (it never is), and there are other power profile changes between them, but it certainly isn't remotely significant of a power draw.
Because in the profile of a smartphone it is absolutely negligible. Your phone is always in radio contact with the cell tower, that absolutely dwarfing all other power consumers. When you turn it actively on, the screen and the CPU absolutely dominate power consumption. There is no case where memory on smartphones is remotely a significant power consumer.
The iPhone has 1GB because that maximizes Apple profits. Every justification are like the hilariously silly claims when the iPhone was 3.5" so many had to justify why 3.5" was the ultimate size and aspect ratio. And it'll immediately shift again once Apple adds 2GB and a 5" screen to the iPhone 6.
The question is, why doesn't Apple try to earn even more money by keeping the same A6 processor as in iPhone 5? Why bother to upgrade to A7 at all? Also, including an extra GB of RAM would cost Apple much less than upgrading to a whole new processor, don't you think?
Apple's marketing tends to stick to what users will understand. Saying "Your apps will run twice as fast" people get, "Your phone will have twice as much memory", people don't.
Now, granted, saying 64-bit is somewhat idiosyncratic for Apple, but if you look closely they're saying that it's the first phone with a 64-bit processor. While people probably don't know why they want a 64-bit processor, they do know why they want a phone with cutting edge technology. As for why it's good, they keep referring to speed, which people can relate to. Still slightly unusual though.
It doesn't need it, but it certainly helps, as benchmarks and my own development of an app that does live video effects has shown.
so that they could have a "killer feature" to lean on if they don't manage to figure out a more novel
Apple has never marketed, nor revealed (to my knowledge) what amount of RAM is in an iOS device model. This is always determined later by a 3rd party.
Neither does the Nexus 4. It's a performance optimization that allows applications to stay resident without being freeze-dried, so to speak, helping multitasking. Could the 5s benefit from more memory? Absolutely, if you're jumping between various applications it makes a difference.
In any case, I wish my iPad had 2GB. With iOS 7 it seems to even discard images on richer webpages that are scrolled out of view. It is starting to seem very confined, and it's hard to view the 5s as a longer term option when it is already living in tight bounds (despite the fact that it has an amazing processor).
I disagree. Retina assets are huge. When programming for iOS (mobile in general) the biggest issue is memory usage and running out. Do anything with a lot of graphics and you start to bump into limits.
The biggest change is an inline retain count, which eliminates the need to perform a costly hash table lookup for retain and release operations in the common case. Since those operations are so common in most Objective-C code, this is a big win.As far as I know, there's nothing preventing that from being increased on future hardware or even on the 5S with future OS updates. I believe the CPU itself supports a 48-bit virtual address space.
And increasing just one aligned integer is certainly cheaper than the bit masking the solution here entails (all of which is neatly hidden away in the 'increment of the correct portion' part).
Additional RAM consumption has costs of its own, in terms of cache usage. Adding an extra 8 bytes for every object in the system is not insignificant. Masking and shifting is extremely cheap.
If you've run the benchmarks and can show your approach is better, by all means, please share.
[1]http://techreport.com/news/25338/new-amd-embedded-roadmap-sh...
What ARM calls ARM related periphery is canonical, whether you think it's silly or not.
However the overarching entity is called ARMv8, with the 64-bit state called AArch64 (which can be contrasted with the AArch32 state, which is also a part of ARMv8) and the instruction set is actually called A64.
And why, exactly, should I care?
"It's important to note that ARM64 includes a full 32-bit compatibility mode that allows running normal 32-bit ARM code without any changes and without emulation."
If ARM64 == AArch64 (which is what you decided), then no, this doesn't make sense. ARMv8 includes both AArch64 and optionally AArch32, the former running A64, the latter running traditional ARM. There are ARMv8 designs that actually can't run any traditional 32-bit code.
As for 32-bit compatibility, that has nothing to do with renaming, it's just me being imprecise. It would be equally imprecise if I had used "AArch64". Thank you for pointing it out, though, I've fixed it to say that it is the A7 that includes 32-bit compatibility.
Do you use the terms IA-32e and EM64T, too (both are/were Intel's official names for what people now typically call x64 or x86-64)?
I, myself, had an AMD Opteron machine on which I ran an x86-64 build, with that name for the architecture, of SUSE Linux before Intel released their EM64T.
http://www.amd.com/us/press-releases/Pages/Press_Release_715...
Dated 8/10/2000.
In short, IA-32e was intel's internal implementation name, which was changed to EM64T which was changed again and is now called INTEL64.
x86-64 / x86_64 / x64 are common names for the ISA created by various OSs, INTEL64 and AMD64 are implementations, the distinction is important because the implementations are not identical.
Debian uses "amd64" and "arm64".
Ah, yes, that must be why everyone in the world calls the 64bit x86 instruction set either AMD64 or IA-32E.
Do you really think that if I had called it AArch64 throughout instead of ARM64, I wouldn't have written... whatever it is you're offended by in terms of conflating various parts? You're nuts!