ISA showdown: Is ARM, x86, or MIPS intrinsically more power efficient?
extremetech.com
extremetech.com
I think instruction density has quite some significance here too - x86 opcodes vary between 1 and 15 bytes with 2-3 being average and ARM has Thumb mode where instructions are either 2 or 4 bytes, but all MIPS instructions are 4 bytes. It also has twice as much L1 as most of the ARM and x86 processors, which apparently didn't help it much. Cache consumes power too, and thus I believe small variable-length encodings (like x86) are ultimately better since they allow for better utilisation of cache; the extra complexity in the decoder to handle this, which basically amounts to a few barrel shifters, is almost nothing in comparison to the area and power that more cache would need.
The entire reason CISC architectures emphasized complex multi-cycle instruction execution is because memory accesses were orders of magnitude slower than the processor and data storage was extremely limited.
When considering cache, these points are all true again. There's a common belief about optimising for x86 to avoid the smaller but slower "CISC" instructions, but in situations like tight loops, an instruction that's 2-3x slower individually can be better than the faster longer one(s) if it means the difference between code and data staying in cache or a 10x+ slowdown from a cache miss somewhere else. Especially on an OoO/superscalar design where the slower instruction can be executed in parallel with other nondependent ones. (Intel/AMD's focus on speeding up these small CISC instructions - which they have done - is possibly one of the reasons why x86 performance continues to improve.)
Really the notable thing to me isn't the ISA nonsense at all. It's how singular a success the Cortex A9 core is. It came at exactly the right moment in history and hit exactly the right sweet spot, being significantly beefier than the A8 yet only minimally more power-hungry. Krait has followed on pretty well, but the A15 can almost be considered a failure at this point.
It would be interesting to see how the A7 also compares to the A8/A9, since it's supposed to be a more power-efficient version of the A15.
Like ARM, MIPS has also 16 bit extensions, two in fact: MIPS16 and microMIPS.
> [cut] thus I believe small variable-length encodings (like x86) are ultimately better since they allow for better utilisation of cache
I disagree: mixed 16/32 bit RISCs ISA have nearly the same code density as x86 and are far simpler to decode.
So beside software compatibility (the killer feature) (well except for Intel of course), the x86 ISA has NO advantage.
My rule of thumb: if something is conceptually "pure" instead of a complicated, carefully balanced mix of grey, its usually not worth considering.
I found this to be true for programming languages, ISAs, politics etc.
It is interesting how purity has a very strong allure - maybe our brains are naturally drawn to a reduced state of complexity, and thus energy consumption?
Or maybe complicated more often than not is just not a "carefully balanced mix of grey" but more of a clusterfuck .. and we learned to be wary of it.
Have a look at this: http://www.infoq.com/presentations/Simple-Made-Easy
If elegance and simplicity are achievable without making too many sacrifices, great! I'd choose Clojure over C++ any day.
http://research.cs.wisc.edu/vertical/papers/2013/isa-power-s...
It seems to me as if this article is mostly linkbait simply by reason of it failing to provide anything more than vague phrases about the source: "This paper is an updated version of one I’ve referenced in previous stories, ... the team from the University of Wisconsin"
Half-baked studies frequently attempt to shout down the real hard science.
" Our methodical investigation demonstrates the role of ISA in modern microprocessors’ performance and energy efficiency. We find that ARM and x86 processors are simply engineering design points optimized for different levels of performance, and there is nothing fundamentally more energy efficient in one ISA class or the other. The ISA being RISC or CISC seems irrelevant."
http://research.cs.wisc.edu/vertical/papers/2013/isa-power-s...
Rough up-to values for various things I can think of / see:
- 10% because they're using an OS rather than running binaries straight.
- 15% because GCC is odd and -O3 does even stranger things, particularly when it comes to energy.
- 15% because their benchmarks are large workloads rather than microbenchmarks that may better target the architecture rather than being huge lumps (would exacerbate GCC strangeness).
- 17% because they're measuring board rather than CPU power supply (the claim that SoC-based ARM development boards cannot have processor power isolated is questionable - I've seen Beagleboards with the CPU power supply isolated)
- 10% because they're measuring energy consumption at a low resolution (their equipment measures in Hz when there's kit that happily measures in kHz or MHz).
Of course, some of these will cancel, and others will be nowhere near as bad as stated. It also doesn't introduce order-of-magnitude changes to the conclusions, although a few of the 'Key Findings' may want questioning.
For implementations more advanced than this... I don't think you can make any such claim based on ISA. x86 may be at a slight disadvantage due to decoding, but that's about it.
To factor out the impact of technology, present technology-independent power by scaling all processors to 45nm and normalizing the frequency to 1 GHz.
The implicit argument of the paper is that Intel could produce a direct size+power+speed replacement for a phone-scale ARM processor, they just need to dial the knobs to small+small+slower. The counter argument is that they have tried but not come close. The Atom line is roughly comparable with respect to speed, but size and power are a problem. The Galileo processor is roughly comparable with respect to power and size but speed is horribly lacking.
The question for Intel is dialing down the profit knob: how much of a hit do they want to take on each unit shipped, by competing with ARM for tiny phone chips.
For those interested, there is a master thesis from the '90 that discussed a prototype reversible ISA + RTL that wasted (in theory) no power for the logic (non-IO) parts [2].
[1] http://en.wikipedia.org/wiki/Reversible_computing [2] http://dspace.mit.edu/bitstream/handle/1721.1/36039/33342527...
I thought these were all RISC processors when you get past the instruction decoder.
Which is how a lot of mainframes were implemented in the 60s/70s (e.g. KL-10).
Power consumption by the instruction decoder vs the power consumption of additional cache&memory bandwidth.
It's amusing that despite the complaints towards X86 in the 90's nowadays it's actually a really good instruction packing format (though it became really sensible only after AMD64).
The original difference between RISC and CISC was that RISC eschewed arithmetic+memory-operations in the same instruction. Both Intel and AMD processors violate this commandment. Instead their decomposition of instructions into uops is based more on the more pragmatic notion of choosing uops that are easy to pipeline and execute out of order.
A couple of citations: http://arxiv.org/pdf/1406.0117v1.pdf (algorithms / data types) http://arxiv.org/pdf/1303.6485.pdf (compiler flags)