Qualcomm Insider: Apple 64-Bit Chip `Hit Us in the Gut'
blog.hubspot.com
blog.hubspot.com
People try to pin their mobile success on Apple fanboys or their "magic" user experience/ecosystem but Apple is simply just x-many months ahead of everyone's product roadmap and continually throwing curveballs at the rest of the players in the market. "Real world" benchmarks aside their 64-bit CPU has everyone else distracted, rethinking their product roadmaps, and demanding it in their devices while Apple works on it's next trick.
b - screen on iphone 5S is already very strong.
c - After lots of thinking I agree with Apple. Current form factor is just right for communication device [1]. If you need bigger screen, chances are you do not want a phone - you want a tablet.
[1] Iphone user here, take this with a grain of salt. Over 3 models of iphone I never had a case. Phone is small enough for a proper grip with one hand while using thumb to control it.
Anecdotal experience: I know 12 people with phones that have large screens. 7 of them had screens cracked due to dropping the phone. Every single one has a case (People who dropped their phones got a case after). What is the point of getting super light and thin phone just to put bulky case on it?
All that thinking got you one vote - unfortunately for apple, most votes want a bigger phone (the numbers don't lie).
>If you need bigger screen, chances are you do not want a phone - you want a tablet.
So now people buying larger phones don't really know what they want? This is typical, flawed Apple superiority complex thinking.
Who wants to carry a phone AND a tablet around (one in each pocket?) How is this even portable or feasible in most cases? No one wants to carry a tablet around with them everywhere they go. A larger phone is a logical compromise. Besides, economics is often the biggest decider. If someone really wants a tablet, but really needs a phone are they going to buy an iPhone AND an iPad? Can they even afford that? or can they really justify spending $1,000 for it when they can spend less than half of that for a well polished Android phablet?
This kind of remarks is really annoying because it suggests that you know better than us (the people who want larger screens) when you really don't. Either that, or you are incredibly blind to the difference between a freaking tablet (7", designed for browsing on WiFi) and a telephone with a decently sized screen (4.7 ~4.8" compared to the 4" of the iPhone - this makes a HUGE difference).
Yes, 20% of all users is still a lot of users, but it's not the 80% that Apple has traditionally targeted.
Who are you to decide that for others? When I pick up an iPhone, the screen is way too small. I have a 5.2" display on my phone, and so 4" seems comically small.
And I also have a Nexus 7, so I can assure you that there's a huge difference between a 5" phone and a 7" tablet. The latter cannot substitute for the former.
> Anecdotal experience: I know 12 people with phones that have large screens. 7 of them had screens cracked due to dropping the phone
And in my anecdotal experience, everyone I know with a broken screen has an iPhone. So this sort of conjecturing isn't very helpful.
I don't know. I think the lack of water resistance makes the iphone obsolete for many.
Apple deserves credit for pushing a more aggressive roadmap than most, and I'll be buying another round of systems here shortly, but being 'first' isn't enough. The Transmeta chip in the MM10 was dual-core 128 bit VLIW with a soft x-86 translation layer. The next chip was 256 bit. Both were very low power devices for their day.
Except these performance gains most likely would of been seen on a 32 bit version of the chip.
It's interesting to see how much marketing is put around 32 vs 64 bit chips and the effect it has on how people market. The biggest benefit is the ability to access more than 4GB of memory. For reference, the iPhone 5S has 1GB of memory. Eliminating the need for PAE and other expensive workarounds is definitely a bug plus. i64 and cheap doubles are certainly a bonus too.
I suspect the primary reason for moving to 64 bit is not for performance gains now, but for future compatibility. Eventually, the iPhone 5S will be the oldest supported iPhone and running a 64 bit processor will reduce maintenance costs of needing to maintain a 32 and 64 bit line of software.
I think this would be a much more informative article if the author had interviewed a few experts to discuss the benefits of 64 bits over 32 bits rather than just say they go faster which at this point is debatable.
It's because it's a much more approachable pitch than "we doubled the size of the reorder buffers and added another instruction dispatch port." In Apple's case, they're simply using "64-bit" as a proxy label for all the other architectural improvements in Cyclone.
No. The biggest benefit as having an addressable virtual address space which is in the tera/peta/exa byte range. Which makes a lot of things simple and efficient, even if you don't have a lot of physical memory. Read about e.g. LMDB's design or Varnish's design for good examples.
The other significant benefit is more registers.
It was badly worded on my part, I was in a hurry. :)
<quote>
Conclusion The "64-bit" A7 is not just a marketing gimmic, but neither is it an amazing breakthrough that enables a new class of applications. The truth, as happens often, lies in between.
The simple fact of moving to 64-bit does little. It makes for slightly faster computations in some cases, somewhat higher memory usage for most programs, and makes certain programming techniques more viable. Overall, it's not hugely significant.
The ARM architecture changed a bunch of other things in its transition to 64-bit. An increased number of registers and a revised, streamlined instruction set make for a nice performance gain over 32-bit ARM.
Apple took advantage of the transition to make some changes of their own. The biggest change is an inline retain count, which eliminates the need to perform a costly hash table lookup for retain and release operations in the common case. Since those operations are so common in most Objective-C code, this is a big win. Per-object resource cleanup flags make object deallocation quite a bit faster in certain cases. All in all, the cost of creating and destroying an object is roughly cut in half. Tagged pointers also make for a nice performance win as well as reduced memory use.
</quote>
https://www.mikeash.com/pyblog/friday-qa-2013-09-27-arm64-an...
Android's Java runtime is going to need a lot of work to make it pull off a runtime feat like tagged pointers.
http://anandtech.com/show/7335/the-iphone-5s-review/4
Basically, the new arch increased registers, increased register widths, and cleaned up some architectural warts in Arm that hurt performance.
The "4GB memory" issue is a red herring -- 64-bit is about making all the paths through the chip wider to process more per cycle.
All those companies are panicking, not because they are being left behind technologically, but they are being left behind in the buzzword war.
The simple fact of moving to 64-bit does little. It makes for slightly faster computations in some cases, somewhat higher memory usage for most programs, and makes certain programming techniques more viable. Overall, it's not hugely significant.
The ARM architecture changed a bunch of other things in its transition to 64-bit. An increased number of registers and a revised, streamlined instruction set make for a nice performance gain over 32-bit ARM.
Apple took advantage of the transition to make some changes of their own. The biggest change is an inline retain count, which eliminates the need to perform a costly hash table lookup for retain and release operations in the common case. Since those operations are so common in most Objective-C code, this is a big win. Per-object resource cleanup flags make object deallocation quite a bit faster in certain cases. All in all, the cost of creating and destroying an object is roughly cut in half. Tagged pointers also make for a nice performance win as well as reduced memory use."
[1] https://www.mikeash.com/pyblog/friday-qa-2013-09-27-arm64-an...
The transition could have happen later, maybe, with not too much repercussion, but why not make the jump now, especially when you have a whole new OS engineered to take advantage of it?
Processor word width has absolutely nothing to do with memory size. e.g. 8-bit processors regularly interface with 16-bit-wide memory buses, and 64-bit processors often allow use of 32-bit-wide pointers.
What the word width does affect is how much data can be processed by each instruction. Certain algorithms benefit immensely from this (you can easily double throughput), but the code needs to be written to take advantage of this (or at least, not to squander the advantage), and needs to be compiled to use 64-bit-wide registers.
Huh? How would you address high positions on your 8Gb memory with a 32 bits pointer that will overflow after 4Gb?
If the processor does have special pointer registers (e.g. 6502, 6809, 8086, AVR, etc.), you use those.
"Adding it all together, it’s a pretty big win. My casual benchmarking indicates that basic object creation and destruction takes about 380ns on a 5S running in 32-bit mode, while it’s only about 200ns when running in 64-bit mode. If any instance of the class has ever had a weak reference and an associated object set, the 32-bit time rises to about 480ns, while the 64-bit time remains around 200ns for any instances that were not themselves the target.
In short, the improvements to Apple’s runtime make it so that object allocation in 64-bit mode costs only 40-50% of what it does in 32-bit mode. If your app creates and destroys a lot of objects, that’s a big deal."
https://www.mikeash.com/pyblog/friday-qa-2013-09-27-arm64-an...
That's what happened on the desktop. I think Apple only shipped one generation of 32-bit Intel processors. Apple was shipping 64-bit CPUs for years when many Wintel computers were still on 32. So when it came time to say 'We're going 64-bit only' with Mavericks Apple left almost no-one behind. On the other hand Microsoft still supports 32-bit Windows 8.
Apple is preparing to make their lives easier in the future. For today's user, it's usefulness may not be too much higher than the level of a buzzword.
We're less than one iteration of Moore's Law away from the 32-bit address space not being enough, even if you're just concerned about memory sizes of mobile devices. Getting this done doesn't seem particularly premature.
1. Think about designs like e.g. the Varnish cache versus squid where the latter team has poured an enormous amount of effort into features which allowed them to manage storage manually, work which is increasingly counterproductive when you can simply mmap() 4GB and let the kernel perform that complex, error-prone task for you while you focus on more interesting features.
For Apple the story is much more than the 64-bit A7 chip, but iOS 7 and Xcode all being released with highly polished support for the A7.
So how do I figure this? Because their first 64-bit chip they've just announced a few days ago (Snapdragon 410) is an ARM one! That blew my mind. Why? Because Qualcomm licenses ARM's architecture in order to build its own cores, and that should mean that they are coming out with a next-gen core faster than ARM's "stock cores". And yet their very first 64-bit chip is based on ARM's stock cores?
That made it pretty obvious that Qualcomm was not prepared at all to make an ARMv8 chip, and I think it was mainly laziness, because of the nice position they enjoy in the Android market right now. They thought they could squeeze 3 years out of the 32-bit Krait, when normally it should've been two.
Well the joke is on them. I certainly won't buy any phone in 2014 that isn't 64-bit, since I usually change my phones every 2-3 years, and I don't want to buy one of the very last 32-bit phones.
I was actually expecting Qualcomm to release its own custom-made ARMv8 core early 2014 - and on 20nm! So I can't believe Qualcomm was actually going to rest on its laurels to such degree that they weren't even planning to switch to 64-bit, and probably 20nm either, even in 2014! That's insane, and only a company that thinks it has too much monopoly power in a market would do that.
As Andy Grove used to say - stay paranoid about your competition. A paranoid company would've expected at least to some degree Apple to release its first 64-bit chip this year. But they didn't even have to do that. They just had to follow the original plan. The one that says release a new core every 24 months, and on a new process node (the one ARM is also following). What ever happened to that plan?
It took years for the PC market to transition from 32 to 64, and it'll take years for mobile to do the same. The amount of time it will take means competitors shouldn't worry because by the time 64 bit processors make sense, everyone will be manufacturing one. Sure memory addressing is just part of the picture and a 64 bit processor means data bandwidth is double that of a 32 bit chip which can be seen below 4gb of RAM, I think at present the iPhone isn't taking advantage of such improvements.
Programs compiled by a compiler aware of those distictions should significantly, measurably, perceptibly outperform those compiled for the older more constrained architectures.
http://internetzona.pl/wp-content/uploads/2013/03/spfsg.jpg
The S4 still fits in your pocket. It's 2.75 inches wide. iPhone 5 is 2.31 inches wide.
Anyways, I realize some ppl like the current form factor, I just wish Apple would give people the choice. I know in my circle of friends they're losing a lot of business to Android just because of the screen size.
Except that's just FUD: http://www.phonearena.com/news/New-Android-and-iOS-fragmenta...
I don't follow the mobile space, so I'm probably missing something.
https://mikeash.com/pyblog/friday-qa-2013-09-27-arm64-and-yo...
How about a device with "physical memory of 1 or 2GB and tight power constraints"?
To quote the article, "the new 64-bit A7 chip is the smartphone equivalent of a big V12 engine." Sure, but there's a reason why you don't drive a car with a V12 engine, too.
Fortunately, phones are not cars, and engineers don't make design decisions on the basis of strained or completely bogus metaphors.
So as far as "do 64-bit processors have tradeoffs that are not worth it in current cell phones", I think it's still a fair question.
It's been said many other times elsewhere in the comments, but to summarize:
1) One of the biggest benefits of a 64-bit processor is the ability to address a gigantic memory space. Phones just don't have that much memory today.
2) Even if they did, the types of processes that are run on phones usually aren't the kind that need that much memory. "Isn't that just 640k is enough for everyone"? Sure, but memory isn't free--if nothing else, a larger DRAM burns more power. So if you want to run a 8GB SQL server on your phone, it's going to cost you in terms of battery life.
3) Another big downside of a 64-bit process is that pointers are twice as large. So your 64-bit process actually consumes more memory than a 32-bit process; in programs with pointer-heavy data structures (like, say, trees) this can result in the 32-bit process being able to solve bigger problems than the 32-bit process. I have seen this in the embedded world when we moved a piece of network equipment from a 32-bit processor to a 64-bit processor. We doubled the memory but that only increased the solvable problem size by about 1.4.
4) Additionally, 64-bit process will tend to blow out your caches a lot faster.
5) If you've got a large data set that you'd like to plow through "big chunk" at a time, going to 64-bits at a time might be a win. But a 32-bit processor with vector instructions is usually quite good, too.
6) Does that 64-bit processor have more transistors? Probably. Does that mean it burns more power? Probably.
There are a lot of good things about that 64-bit processor that I am not mentioning, but the point is they are not free. A lot of these tradeoffs are hard to measure by outsiders because they didn't just change one variable: in this case, the move to 64-bit introduced new instructions, more registers, was concurrent with a smaller and faster transistor size, bigger caches, etc, etc.
It's important to consider what you could have done otherwise. Possibly you could have had a 32-bit processor with more cores. Possibly you could have had similar performance with more battery life. Possibly you could have had a lot of things, but if you just think "64 is bigger than 32 and therefore better" you will miss those possibilities.
This is a good point, and it's the first time I'm hearing it stated so succinctly.
With 64 bit mainstream chips, the manufacturers got to roll them out slowly, first to the scientific computing community and then into server applications before finally tackling widespread desktop adoption. Given that there's not an equivalent path in the mobile space, I guess you have to test the waters somewhere.
For example adding 64-bit numbers; finding the end of zero-terminated strings; certain hashing and encryption operations; and so on.
It's also good marketing, of course.
14 registers is pretty tight. 31 registers are better, and doubling the width helps for structure locals and parameters (which Clang/LLVM fortunately does a good job of keeping in registers). (I do a lot of work on a processor 60-odd 64-bit registers, and even then GCC decides to spill registers now and then.)
And whether Apple wants to buy into the chip game, well, good luck with that.
All 64-bit does at the mobile level is introduce a ton of power loss for no real performance gain. The power usage Apple is reporting is likely due to other technological improvements that more than compensate for 64-bit's loss.
The first poster had it right. All this does is screw up the rest of the industry, make them scramble, while Apple can buy time to figure out what to do next.
Side note, Qualcomm's layoffs have little to do with their processor line...
As for addressing 128 bits of memory: that's more than a century off even if memory continues to double every two years (which doesn't seem likely to begin with). It's actually plausible that the step to 64-bit addressing was the last one, ever.
In most code settings, this pollutes the cache (you can only hold half as many pointers), and leads to slightly weaker performance.
64-bit computing is NOT an advantage in the phone world. It is a massive advantage for database applications or large web services... but certainly not for phone apps in the near future.
(And while one could argue that a vector register isn't a "real" general-purpose register, the primary difference is that you can't perform fullwidth arithmetic, which is of very little use > 64 bits anyway.)
No seriously, there is little to no speed difference in 32-bit vs 64-bit computing. The REAL benefit is the ability to address memory beyond 4GBs. But even the most high-end phones are stuck with 2GB... hell, a number of laptops are still shipping with 2GB of RAM. Let alone phones. The 64-bit "advantage" is almost entirely a marketing gimmick.
That said, it is a known fact that Apple's iPhone chip is leagues ahead of its competitors in terms of performance / watt. Qualcomm has a value buy with Snapdragon and their integrated LTE chip + FCC conformance. But Apple has full integration, and controls the software / hardware from bottom up. That is a real advantage that is leading to improved battery lives and faster performance.
(This is less an "Apple Advantage" and more of an "Android disadvantage". Android should be able to catch up if they got their act together... but the reliance on Dalvik-VM code and unoptimized APIs leads to noticeably worse battery life)
Beyond that, Apple is beginning to invest into state-of-the-art foundries, possibly to create chips in house by the year 2016 or 2017. They've also bought out the high-end 20nm wafers through 2014, forcing their competitors to lag behind on the older 28nm process nodes. (Hell, even AMD / NVidia are feeling the sting. All AMD / NVidia roadmaps to better GPUs are at 28nm technology... Only Intel: who has reached 14nm on their in-house labs, remains unaffected by Apple's purchasing power)
When your company has $100 Billion in cash, you can afford to have a process node advantage over your opponents.
Very interesting. Would you have a source? I know Apple did the same "trick" with touch screens when the iPhone was introduced. I think these actions are typical of Cook - from what I understand he was earlier (before he became CEO) responsible for the whole inventory chain management and it's likely an area he's still involved in.
http://semiaccurate.com/2013/07/12/apple-has-their-own-fab/ http://www.tomshardware.com/news/TSMC-Samsung-Apple-UMC-A-Se... http://appleinsider.com/articles/13/07/12/rumor-apple-buys-i...
On Apple eating up wafer supply: its mostly speculation from even more illegitimate rumor sites. But those keeping up with the current roadmaps note that Apple and Samsung are reaching 20nm far before their competitors... and even AMD / NVidia have their 20nm plans pushed out to 2015 or later.
"Proof" is nonexistent, but the product roadmaps I've seen seem to match the rumors.
I'm not familiar enough with the particulars of ARM to answer confidently for floating point operations, but to take an example that's not usually special-cased, say bit vector arithmetic, yes, those operations will execute twice as quickly if they are vectored.
On x86 though, both 32-bit and 64-bit did double precision vectors just fine, so it didn't really apply there (except that the fp register count was doubled).
Most SIMD code is heavy number-crunching stuff like multimedia or GPU shaders. But much of that low-level handling is handled off CPU on phone platforms. It is simply more power efficient to have a hardware decoder of multimedia.
GCC has supported autovectorization for a while now.
"Unoptimized code will be slow" isn't a great argument anyway. There's not much a processor can do to help that.
Besides, most ARM chips supported vectorized code anyway. You know, NEON? http://www.arm.com/products/processors/technologies/neon.php
ARM64 does NOT grant you the vectorized instruction advantage. Qualcomm Snapdragon Krait have supported NEON for some time already.
http://www.anandtech.com/show/5559/qualcomm-snapdragon-s4-kr...
So many people are arguing here, but clearly few of you people have even worked with ARM chips at the assembly level.
(1) it's enabled by default at -O3; (2) the loop constructs needed are fairly simple; (3) arguing that unoptimized code will be slow is still a poor argument.
ARM64 does NOT grant you the vectorized instruction advantage.
32-bit ARM NEON does not support vectorized doubles. 64-bit ARM NEON does. Source: http://en.wikipedia.org/wiki/ARM_NEON#Advanced_SIMD_.28NEON....
So many people are arguing here, but clearly few of you people have even worked with ARM chips at the assembly level.
Yep. Thankfully I can back up my arguments with quoted facts.
EDIT: and I already granted that vectorized and floating-point operations don't necessarily benefit from larger register widths, so I don't know why you're even arguing. Let alone the OP wasn't even asking specifically about ARM!
No, there really is a huge performance gain for 64-bit processing for certain algorithms when coded correctly. Basically anything that works with vector-like data can easily benefit. I'm sure there are lots of mobile multimedia developers who relish the change to 64-bit.
(Of course it's possible the 32-bit predecessor to this chip special-cased certain 64-bit operations, e.g. double float arithmetic, in which case even fewer algorithms would benefit from widening registers across the board. I'm not familiar enough with ARM architecture to comment on this.)
Game Programmers will prefer a faster GPU, since none of that stuff is actually calculated on CPUs now-a-days. (in fact, Apple's superior GPU is one of the reasons why it "feels" so much faster than many Android stuff).
So unless you're gonna be doing software-decode of H265 (or some other future codec), or something... my bet is that multi-media processing will remain the same. It will go to the dedicated multimedia DSP that is on every phone, and be translated extremely efficiently (powerwise).
Uh, no? Yes, video decode for common formats is hardware-accelerated, but I've never seen dedicated Fourier transform hardware in consumer hardware, and I can't think of any other "and the like" algorithms that are hardware accelerated not at a CPU register level.
Game Programmers will prefer a faster GPU, since none of that stuff is actually calculated on CPUs now-a-days
Mm, I think this is dubious. I agree, GPUs are better than CPUs for many multimedia applications, but getting data to and from GPUs is not fast. And of all the multimedia applications I run on my desktop (mplayer, Audacity, the Gimp, Inkscape), none currently use the GPU except for maybe mplayer for certain videos.
In fact, Intel's DxVA implementation explicitly has an iDCT accelerator. See this paper for details: http://download-software.intel.com/sites/default/files/artic...
I assume a lot of people watch Youtube on Windows computers, amirite? The iDCT is basically a Fourier Transform as far as the math is concerned. Other portions of the H264 codec (such as motion compensation) are similarly increasingly hardware-accelerated... even on crappy integrated GPUs like the old GMA950.
Phone hardware on the other hand, is basically state-of-the-art. I wouldn't be surprised if phones of today had superior hardware decoders than the crap that Intel churnned out for the bottom-barrel consumers back in 2009.
Unfortunately you entirely missed my point about everything other than video decoding. Bandwidth between the CPU and GPU quickly becomes the bottleneck, unless you're able to move most of your processing onto the GPU, which I granted you was the right thing to do. But also as I stated, none of the popular software I use actually does this. It is all optimized for CPU processing.
DxVA suffers from this same issue, i.e. you have to be very careful around moving data to & from the GPU: http://en.wikipedia.org/wiki/DXVA#DXVA2_implementations:_nat...
EDIT: And in case you think I'm talking out of my ass, I work on a high-performance embedded product. We recently switched from a 32-bit to a 64-bit version of the (ARM-like) processor we use. Nearly every single one of our major algorithms benefited from the increased register width (although we did have to slightly modify some of them to do so). And we don't even use multimedia operations. A lot of the gains come from simply moving less stuff around, which, when you have to process a packet every 40 cycles, really adds up.
If it is faster, it's not because it 'handles data in bigger chunks'.
The whole industry will be naked in a week.
All of this does not really matter anyway. iPhone has always had better performance than almost every android phone. This has not saved them from losing massive market share. The PC/Mac history is repeating with Android/IOS.
Exactly: Apple dominates the high end, and is at 0% of the low (some might say junk) end.
Where did you get the idea that Apple dominates at the high end? They haven't for a while now.
The Galaxy S4 was outselling the iPhone 5 all by itself for several months (before the 5s came out). The collective of high end Androids have taken over the high end market on a seemingly permanent basis. Apple is at less than 50% and is on a steady decline vs the (high end) Android collective of phones.
sounds legit to me.