ARM details its new high-end CPU core, Cortex A72
arstechnica.com
arstechnica.com
For a long time Intel has been top dog in the "instructions per clock" or IPC space. So if you wanted performance you used their chips, except you paid for that by consuming a lot of power. ARM on the other hand has always tried to be the 'low power' chip which you could embed and run on batteries, and ultimately slower than the Intel architecture.
But into this a couple of interesting market realities intruded, the most obvious was that at some point computers were "fast enough" for enough people to make a durable market for lower power machines. That really took off course in the smart phone and tablet market, but is inching into the "low end" server market.
If ARM can get better at performance faster than Intel can get better at "low power" they can really put a dent into Intels market dominance in the larger computer space. And as a competitor with a completely different ISA they limit Intels ability to compete with lawsuits and/or changes to the "standard."
... That's not how it works. You don't add front-end and back-end width to get issue width. I know that Ars' architectural chops took a hit when Stokes left, but this is ridiculous.
You take the minimum of the two[1], except that (per AnandTech's far better article [2]) the decoder is actually three fused macro-ops wide, not instructions, so the truth is actually somewhere between three and five instructions wide, depending on workload.
[1] or maybe just the back-end width, since most modern CPUs can short-circuit decode in tight loops.
[2] http://www.anandtech.com/show/9184/arm-reveals-cortex-a72-ar...
It's like when you add a few cortex m0's to your phone, and now you have a quad-core phone.
(It's 8 issue. You can issue 8 instructions to it, it just doesn't process them all in a single cycle :P)
The front end feeds the back end.
3-issue FE means you can give the BE 3 things per cycle.
5-issue BE means you can crunch 5 things from the FE per cycle.
FE = branch prediction, decode, instruction cache
BE = schedule, execute, load/store, data cache
I'll leave my comment though for the edification of others.
So ...
Ins 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 ...
FE: [ ] [ ] [Ins0]
BE: [ ] [ ] [ ] [ ] [ ]
FE: [ ] [ ] [Ins1]
BE: [ ] [ ] [ ] [ ] [Ins0]
FE: [ ] [Ins1] [Ins2]
BE: [ ] [ ] [ ] [ ] [Ins0]
FE: [Ins1] [Ins5] [Ins6]
BE: [ ] [Ins3] [Ins4] [Ins2] [Ins0]
FE: [Ins1] [Ins5] [Ins7]
BE: [Ins6] [Ins3] [Ins4] [Ins2] [Ins0]
And then maybe zero retires, that lets Ins1 proceed and it retires 1, 2, 3, 4, which lets 5 proceed, that lets 6 retire and then 7 proceeds.But until I get a closer look at the TRM this is all speculation.
Almost certainly yes. It's supposed to be fully out-of-order, which means it should have a fully functional scheduler in between the FE & BE.
Not to mention, given modern memory latencies (vs. clock speed), letting the FE run ahead is important for performance.
After all, the Cortex-A15 and A57 are also 8-wide issue. Incidentally all 3 are also 3-wide decode and dispatch, but A72 apparently decodes instructions to fewer uops.
IIRC ARM added a 32-entry loop buffer with Cortex-A15. Not sure about later designs.
I've seen people run reddit bots using a Raspberry pi. After a few months its cheaper then a VPS!
I think they're trying to beat both to the punch. :(
Sure, HPC is all about FLOPS, and desktops tend to use quite a bit for presentation nowadays, but I would consider memory bandwidth, I/O bandwidth, thread scaling, and single-thread performance non-FP metrics such as branch latency to be much more important in the general server market.
And who knoww, maybe later we'll see other countries shift to that kind of strategy.
Then came the PC clones and most nations stopped making their own stuff and just imported parts, or whole computers, from Asia.
Even if it remains slightly behind on that metric, ARM is is basically untouchable by Intel in the price/performance metric. Many people, including the author seem to say stuff like "the next Core-M generation will totally beat this anyway!"
Except, Core M is nowhere near the price/performance of Cortex A72. That makes Core M uncompetitive in whatever markets the two would have to co-exist (which they don't right now, because Core M is only used in $1,000+ Windows or Mac OS X ultrabooks).
Even Atom can't be competitive on price/performance. The real (unsubsidized) gap between the two is somewhere around 2x and has remained that way since Atom's existence. The only way Atom can pretend to be competitive on price in the mobile market is because it's subsidized by Intel while it overcharges for the new Atom-based Celerons and Pentiums (Intel charges for them as much as it used to charge for the Haswell versions, and what it would've charged this year for the Broadwell versions - but those don't exist anymore).
The idea that "the slow cores can just handle the background stuff" makes little to no sense in practice.
The most power effective thing to do is not add more cores (and all that entails), but get done the work something wants done as fast as humanly possible so you can go back to sleep. Having slow cores handle this for you sounds great in theory. In practice, it's been a huge bust (The one exception being dedicated-function coprocessors, etc)
This is one reason that, for example, higher speed wifi/etc chipsets often become less power intensive than the old chipsets (past the initial spike from it being new hardware) - they can sleep more
What I've heard suggests that the tradeoffs behind high performance ARM lead to significantly worse power efficiency to the benefit of peak performance; which means that a core like the A53 ends being a very attractive proposition for power saving.
Certainly I don't know why ARM would recommend their big.LITTLE architecture (or at least why people would use it) unless it did.
Furthermore, Amdahl's law means that even in theory it should be advantageous for performance to include a few extra low power cores to boost peak performance (above adding a larger high power core).
Because they want to compete with others who keep upping the core numbers.
"Furthermore, Amdahl's law means that even in theory it should be advantageous for performance to include a few extra low power cores to boost peak performance (above adding a larger high power core)."
Again, in practice, exactly zero of the chips that have done this and been put in phones (nvidia's, etc) have actually saved power by doing so.
Or did things get miraculously better when nobody was looking?
And what if you can actually scale your CPU-bound tasks linearly? Then you can use the extra cores to get the work done faster so you can go back to sleep faster.
I'm not speculating here: I've seen power measurements demonstrating this in real world applications on mobile (in my case, the browser engine). If you can use all your cores to get the work done faster, you can decrease overall power consumption.
You're assuming we aren't going to parallelize our applications. I don't make that assumption.
I'm assuming parallelizing doesn't help enough :) Which it hasn't so far. So let me rephrase "so far, it makes no sense".
and it's not like they haven't existed for quite a while, so it's not "well, it's a chicken and egg problem".
Not according to my numbers. And yes, we have measured the impact of parallel rendering specifically on power consumption on mobile. On real-world Web sites, not random microbenchmarks.
Developers barely know what to do with 4 cores, especially on mobile. The usual outcome is that foreground apps use a core, maybe two if they have async rendering, and background apps can run as well.
Mobile OpenCL is the closest we'll come to massively parallel programming in people's pockets for the foreseeable future.
You just rewrite the popular libraries and game engines. Then lots of applications start to achieve parallel speedups. The point of Servo is in fact to achieve just that—don't forget that a browser is not just an application but a platform, running applications that happen to be written in (up-to-now) essentially-single-threaded JavaScript, HTML, and CSS. But as a specification, CSS and HTML (especially CSS) are in fact quite surprisingly parallelizable, and our tests have shown that existing apps get great speedups unmodified in Servo.
Taking the underlying framework and parallelizing it is an effective technique to take what were previously single-threaded applications and making them parallel. In fact, in a sense that's what superscalar CPUs have already been doing for decades.
For example snoopy caches would I think stop scaling by that point, and in order to fit the cores onto the chip you'd probably have to get rid of a lot of the control flow logic (superscalar, branch predictor), so you'd lose a tonne of performance on singlethreaded code.
You really don't want to move onto 16 cores unless you have to. More likely is more powerful instructions (perhaps true vector instructions).
That seems like vertical scaling instead of horizontal scaling. I'm not sure how much money there is in the former market—maybe look at mainframe markets.
Horizontal scaling is a much easier sell, especially for data centers. I will put money on the fact that Facebook will eventually invest heavily in this architecture design.
What? I'm guessing it is a typo for a 2/4 way set associative cache, or something I know not what.