And while instruction predication does let you avoid a lot of branching nearly for free in the ARM architecture it makes out of order logic much less effective because now there are a lot more dependencies between instructions. That was the reason the Alpha team decided not to use predication, even though they were familiar with the idea. Its also why the Itanium, an in order architecture, used predication.
Consumer electronics is driven by volume & price, so all we are really seeing is small, low power cores becoming acceptably fast for mainstream computing.
All that is really pretty awesome if you're designing your processor to go in order, but if you want to try out of order execution it becomes more complicated. Now you have to check every instruction to see if its predicated, and if it is you now have a new dependency on the previous arithmetic instruction even if there isn't any data dependency between the two. Of course you could argue that if you find a predicated instruction in ARM code then its replacing a branch that would be in x86 code, and that overall your job is no harder. And you could go back and forth arguing over it.
The important thing to remember, though, is the things that give the ARM ISA inherent advantages when you're making low power, in order processors aren't necessarily advantages when you're talking about high performance, high power processors.
http://www.arm.com/products/processors/cortex-a/cortex-a15.p...
A typical improvements for microarchitecture is a factor of 2, spread between different units. For example, we could speed up full word adder (most common roadblock for higher clock frequency in current CPU's) by 10 percents, add scoreboard that allow us to issue 1.3 instructions in average load, add bypass logic that allows us to cut delay by 1/5 (5 cycles for pipeline) and we get 1.11.31.2=1.72 of our previous speed.
I assume that ARM already have all those fancy bypasses and scoreboards. So how would they get that 5 times speed up?
I think it will be true for some special cases, like Javascript interpreter or video decoding. I bet on JS.
It is hard to speed up a mature architecture by five times.
It seems to me that in order for an ARM processor to be viable for something like the MacBook Air (which is already using all the easy power-reduction measures like LED backlights and SSDs), the ARM chip will have to be twice as fast and draw half as much power as an Atom.
x86 has always been predicted to fall behind competing architectures, and it's always kept up--at least for personal computers--because of the large vested interest in keeping all that x86 code running. History is littered with better-architected CPUs that couldn't beat x86. ARM survived because ARM is an embedded processor, and PPC survives as an embedded processor, but Apple's been down the road of trying to shoehorn an embedded processor design into Macs before, and ended up migrating to x86.
One of the factors that help people move across to OSX seems to be the knowledge that they can still install Windows if they need/desire. Perhaps Apple will buy Parallels and include it as part of the OS (although VMWare will have a case if they see this as being anti-competitive)?