AMD Unleashes First-Ever 5 GHz Processor
amd.com
amd.com
This product release is kind of an attempt to show that AMD can actually deliver on their planned strategy.
As for why AMD is going for this route, rather than trying to beat Intel in per clock efficiency? Probably because AMD's resources are severely limited compared to Intel, and this approach offered lower risk at lower cost.
Well, because clock speeds are something they can improve now, and not in $n years time when their next major microarchitecture is ready. Intel had the exact same problem with the Pentium 4, and they were similarly stuck with minor tweaks and desperately increasing clock rates for years before Core was ready.
It was my understanding that P-IV was originally conceived as an architecture that could be clocked up and up for years to come. Developing a new architecture is expensive and risky (see P-IV, Itanium), so the hope was to design something that would scale up so well as manufacturing improved that a few generations of architectures could be skipped, so to speak.
They had hoped the P-IV would eventually reach 10GHz or so. Which made it OK that P-IV retried fewer instructions per clock than the P-III that came before. Scaling up like that isn't such a radical idea; the P6/i686 architecture behind the Pentium Pro, Pentium II and Pentium III had spanned a spectrum from 150MHz to 1.4GHz, after all, nearly an order of magnitude.
But it turned out that somewhere between 3 and 4 GHz, things got really difficult.
"Minor tweaks and desperately increasing clock rates" was more or less the P-IV plan from the get go. It just turned out not to work.
AMD's architects called this pursuit a low gate count
per pipeline stage design. By reducing the number of
gates per pipeline stage, you reduce the time spent in
each stage and can increase the overall frequency of
the processor. If this sounds familiar, it's because
Intel used similar logic in the creation of the Pentium 4.
http://www.anandtech.com/show/4955/the-bulldozer-review-amd-...And since I read this, I really wondered what trick AMD has up its sleeves.
10 years ago AMD had more efficient architecture, Intel - more GHz (remember 2.2 Athlons having 3700 "PR-rating"?). Intel's approach was commercially more successful - customers were still buying GHz, and AMD was trying to educate market about real performance while radiating impression of a looser who just can't get good process and fabs. AMD gave up and decided to pursue P4-like approach for their new architecture, while Intel hit GHz "sound barrier" and went efficiency way by resurrecting PIII style architecture which resulted in Core CPUs. AMD made a huge, strategic mistake 10 years ago. How the execs in charge of Bulldozer have been BS-ing their way inside AMD last 5 years - that is a typical everyday miracle of a big company internal life.
Keeping AMD out of Dell and HP is what kept them out of corporate America. Corporations literally buy PCs by the pallet, and then they pass on the volume discount to employees who want to buy one for home.
Really? I don't see the use case for ordering a PC this way for the home. When I'm buying a home PC, or recommending one for others, it's either:
(a) A bottom-of-the-barrel PC. As long as it has 1GB of memory and more than one core, you can use it for web browsing, Youtube, email, and word processing. This is what non-techies usually want (but they don't know they want it and may get upsold by good marketing). This is what I want unless I'm planning on running a specific application that requires more.
(b) A powerful PC for gaming. It needs a decent discrete GPU if it's going to play current games. Most office PC's don't have one, unless you work for Pixar.
AFAIK the machines purchased by corporations for general office use are usually middle-of-the-road beasts that cost more than category (a) but don't have the discrete GPU of category (b). I'd be guessing they'd be a waste of money for home use, even after the volume discount.
Bulldozer was always supposed to be a high clock long pipe machine, sacrificing some IPC. See eg http://www.anandtech.com/show/5057/the-bulldozer-aftermath-d...
"Per-clock efficiency" is generally not goal in itself in CPU design, absolute performance and efficiency are. Where AMD has stumbled is getting the clock up, probably partly related to their unfortunate fab situation
The speed-demon strategy has seen successes historically, Pentium 4's fate notwithstanding. See eg the DEC 21164 and the IBM z196.
A lot of people will probably be deceived by the high number, thinking that a higher number automatically always is better.
Virtualiztion host performance on our 8 core bulldozer (esx 5.1, private kbs from vmware to try to help, 32gb ram, rad10 zfs san) was so bad (think p4 era) that we finally tracked down how to force the cpu into only using 4 cores, one per real fp core.
The reality is that there is no mainstream scheduler out that that can efficienty use cores set up like that, especially with the long pipelines. I'm not sure it can't be done, but what improvements have been made have been minimal, or just in an academic/not a real os situation.
That's why Intel ships a compiler, duh.
It is true that the # one thing holding that part back was the raw clock speed (as long as you view it more like a 4 core, 8 thread part ala Intel), but i've gone back to speccing intel - it's just not worth being that much of a ginae pig for a firm thats basically trying to scrape by until the arm64 parts start getting stamped.
Was the issue the shared fp units, or the turbo-core? I wonder if you can disable the turbo-core?
Where this can fall apart is if you're trying to use eight homogenous threads at once and the threads have large working set sizes, such that the second thread causes spill out from the per-module caches. Then you have eight threads contending for L3 bandwidth, or if you're really screwed you fill up the L3 and start to hit main memory.
Out of curiosity, have you tried any of the Abu Dhabi Opterons? They doubled the L3 from 8MB to 2x8MB, which I would expect to help by both keeping you out of main memory and reducing contention by splitting each L3 between half as many cores (assuming you don't get the new twice-as-many-cores models).
http://www.anandtech.com/show/7066/amd-announces-fx9590-and-...
Based on an old review:
http://www.anandtech.com/show/6396/the-vishera-review-amd-fx...
and single thread performance:
http://www.anandtech.com/show/6396/the-vishera-review-amd-fx...
If single threaded scales linearly with turbo frequency (and it looks like it might):
The FX8320 (turbo boost 4.0Ghz) scores 240.7, while the FX8350 (turbo boost 4.2Ghz) scores 252.1:
The difference aligns quite nicely: (240.7/4)4.2~252.74
And for 5Ghz should give about: (240.7/4)5~300.88
This is still lower than intel's i5 3570k (302.2 - Turbo 3.8Ghz) and i7 3770k (312.4 - Turbo 3.9Ghz)
And Haswell has even higher performance:
http://www.anandtech.com/show/7003/the-haswell-review-intel-...
I'm not convinced ~11 pixels per second is within the margin of error (the numbers were from the single threaded povray test) -- but 3.8% difference certainly mean very little in the real world. I'd guess it falls within the bracket that is measurable but ignorable ;-)
Also, the "5Ghz chip" will most certainly be AMDs top of the line model?
But what do you expect from a press release? It's written by marketing trolls, not engineers.
That's "first".
PS: Intel and Nvidea do the same thing. http://www.techarp.com/showarticle.aspx?artno=745&pgno=1
And yes there is a market for this. There are certain workloads that are simply not parallelizable -- they're linear chains of dependencies where the output of the first process goes into the second and so on and each step depends on all N-1 steps.
That is to say, for most strictly serial processes, there's an application where you'll want to run it many times on independent data.
http://www.anandtech.com/show/6396/the-vishera-review-amd-fx...
This shows at least part of what I mean - Vishera (slightly lower clocks than what was just announced) loses to Ivy Bridge by a mile in most single-tread tests, but nearly ties in multithreaded ones.
https://intel-newsroom.jive-mobile.com/#jive-content-item?co...
PS: They will even compare old chips with turbo boost disabled to there new chips with turbo boost enabled. See note 3 http://www.intel.com/content/www/us/en/processors/xeon/xeon-...
The "Note 3", however, certainly sounds fishy, but this seems only loosely related to the question at hand.
Do you really want your processors going full bore 24/7 just to prove they can?
The i5 in the new Mac Air runs at 1.3Ghz if both cores are active, but if a single core is active and the environment isn't too hot and ventilation is working, a single core can hit 2.6Ghz. Which is quite humorous -- you might have much better real life performance simply disabling a core.
EDIT: To clarify, I replied because of the supposition by the parent that this turbo mode is "for a few milliseconds". In actual practice it is usually a very significant contributor to performance on modern chips, and as mentioned can be indefinite in some circumstances. Ergo, dramatically more important than implied.
In this case, however, it's an 8-core chip. Very few current workloads will saturate 8 cores (even on heavily taxed database servers), meaning that there is a good chance there is always thermal availability for individual cores (and thus individual threads) to be run at 5Ghz.
It depends. Some tasks, such as processing incoming HTTP requests and building responses, such as web servers do - are embarrassingly parallelizable. And if you have an architecture that scales horizontally, with enough network bandwidth, you can saturate how many cores you want.
Does this have any implications when designing software? (e.g., do things on a single processor in certain situations because it might be faster?)
The article cited in the Wikipedia entry is dated 08-Apr-2008: http://www.theregister.co.uk/2008/04/08/ibm_595_water/
PDF Redbook reference: http://www.redbooks.ibm.com/redbooks/pdfs/sg248050.pdf
Highest clocking POWER processor offered is the 4.42 Ghz P7+ System p 780: http://www-03.ibm.com/systems/power/hardware/780/specs.html
And yea, POWER 6's design was a bad decision.
A: Wait faster!
I suspect that as more software stops being optimized for ~10ms serial disk I/O with huge caches this will become more common and more and faster cores will be a big(er) deal.
I would like to see review though. And pricing. If it has decent single thread performance and that number of cores with all next gen games being multhithreaded by default it could be a compelling processor if it is in the 4770 price range.
That, and I see next gen games more favoring openCL / GL 4.3 compute shaders to offload all their parallel workloads than to aggressively optimize for greater than 4 core processors. Your returns on moving traditionally CPU bound workloads (per agent logic, path finding, collision detection) to compute class GPUs (when available, with the cpu fallback for now) gives you significantly more returns than optimizing for the CPU.
Also, you can take a 4770k to near 5ghz on air. This part is already pushing the thermal limits of the Bulldozer architecture, AMD is just fabbing them out this high speed because they are floundering in this low-per clock performance rut the entire architecture put them in.
Now, I would point any budget oriented gamer to the 4 or 6 core AMD models around $120 - 130, because since they are all unlocked, you can get real performance gains (but terrible power efficiency) over Intel parts below the 4 core unlocked part they put out each generation. Since they are effectively 2 / 3 module parts, they are well suited for the next gen of GPGPU everything in the engine and let the CPU do control flow.
If you even approach $200, the performance gains from jumping from any non-K part to a 4.8ghz 4570k are huge, and that alone outclasses every AMD cpu for gaming, but does trade blows on some titles with the 8 core parts.
For an example, what if a chip used a 10 GHz clock for distribution, and divided it down to 5 GHz everywhere it was actually used (not that I know of any reason to do such a thing besides marketing). Would it be marketable as a 10 GHz chip? The manufacturer would certainly be in hot water if enthusiasts ever found out...
Even without such contrived scenarios, CPUs get different amounts of stuff done per clock.
Something I keep seeing, even on Slashdot and Hacker News, is the idea that a CPU that has to clock higher for a given performance will use more power. It seems to me that if you've got double the clock, the likely explanation is that half the transistors are switching per clock, and power consumption should be orthogonal to clock/IPC ratio.
If anyone's got any contrary ideas on that, I'd love to hear them. All I can think of is that higher clocks would correlate with longer pipelines, but bulldozer's pipeline isn't even that long.
This is like a dog whistle to the EEs, they're going to get all riled up by programmers with screwdrivers. You can model a stereotypical FET gate as a capacitor, all you're really doing is charging and discharging capacitors either in FET gates or the transmission line theoretical capacitance. Right out of the C=Q/V definition of what capacitance is, mushed up against some ohms law and some algebra, and you end up with P=C times V squared times F. So you can see the intense excitement in lowering core voltages, making gates and lines smaller (lowering C) all in a tradeoff to improve the P/F or F/P (whatever) ratio.
The important part is its pretty easy, right outta ohms law and the def of what capacitance is, power is directly proportional to frequency.
There's also the fact that your transistors have a particular voltage that they switch state at, which means that they switch faster if you drive the gate/line capacitance with a higher voltage.
Which means that chips designed for lower frequencies can be designed to use lower voltages, which can save far more power than what would be directly proportional to the lower frequency.
yes, right out of the equation provided.
In "CS" terms that may be better understood on HN than "EE" terms, electrical power scales O(n squared) with voltage and O(n) with frequency.
If you really wanna get people riled up and talking you can roll out the old power "EE" stuff about maximum power transfer happening when source and sink impedance are the same, and you want to get the most bang for your buck so you'd like that, right, and a transistor gate being near infinite resistance would imply ... Or if you like to think about interconnects being signal to noise level limited, then an RF analysis about noise voltage across a resistor vs preamp noise figure vs current bias from a communications standpoint would imply... But it turns out in practice most of the time, the first mental model is by far the most effective way to look at it compared to these.
Suppose CPU A has an adder, that takes one clock cycle to run an add instruction. When two registers are being added, the instruction goes thru the entire adder in one clock cycle and affects on average some % of the transistors.
Suppose CPU B has a pipelined adder that takes two clock cycles to run an add instruction. When two registers are being added, the instruction goes thru half of the adder in one cycle, and the other half in the next cycle, and affects about half of that same % of the transistors each time. BUT! This is a pipelined adder, and doesn't just do one instruction at a time. During the first cycle, when our instruction is in the first part of the adder, some other add instruction is still going thru the second part of the adder and affecting the other half of whatever % of the transistors. And during the second cycle of our instruction, the next instruction is going thru the first half. So even tho any one instruction only affects half of the adder at a time, the entire adder still gets affected every clock cycle.
Roughly speaking, power used = transistors switching per unit time. Performance should also follow that pretty closely, depending on the efficiency of the design. At some level, you should be able to look at any instruction and find a corresponding number of transistors that need to switch for it to execute.
Deep pipelining keeps more silicon active at any given time, increasing both performance and power consumption. Because of cache misses and the like, efficiency will drop somewhat. Double the stages also doesn't quite equal double the switches per time, for various reasons. Therefore, deeper pipelines = worse performance per watt but better performance per dollar (not sure how well that'll hold in ridiculous cases like Prescott).
From what I heard, Bulldozer only has one more stage than Haswell (15 vs. 14, don't quote me on that) - not nearly enough to account for the differences we see between them.
What I'm noting is that there are many, many more factors at play than just pipelining. In the case of Bulldozer, I've been hearing quite a bit about minor parts that they found needed more work, most notably branch prediction. It sounds like they've got lots of things that will improve performance with no power or die size downsides. The number I saw bandied about for Steamroller was a 30% performance increase. I have some trouble believing it's quite that big, but if they pull it off, that will be an amazing chip for being 32nm. It hints to me that the macroscale architecture is A-OK, and they just screwed up some small but important things.
Nope; a lot of the latches are switching every cycle, so power is higher at higher frequency. This is what doomed NetBurst-style design.
Just making up some numbers, how about 30% of gates switch on every clock, and 3x the switching speed for modern gates (it's probably much higher, but I'm being conservative here):
NetBurst: (0.3 * 3) / ((0.3 * 3) + (0.7 * 6)) = 17.6% power
Bulldozer: (0.3 * 4.5) / ((0.3 * 4.5) + (0.7 * 18)) = 9.7% power
Sandy Bridge: (0.3 * 3.6) / ((0.3 * 3.6) + (0.7 * 18)) = 7.9% power
So basically, NetBurst is ridiculous, though that shouldn't be news to anyone. Bulldozer doesn't look to be doing so bad as all that, and the numbers improve if the speed is more than 3x.
(I have no idea what the real numbers are, if someone tells me I'll update this.)
If we told intel that they could burn up to 350watts on the CPU and a 25lbs heatsink was acceptable, we'd probably have 10ghz processors. Problem is, there isn't a large market for that. Home users don't want a big ugly and noisy box and server buyers would prefer power and heat savings. Supercomputers just tie all this stuff together instead of creating some monster single-core.
Actually, this was the strategy with the pentium 4. It was a fast and power hungry single-core. Turns out, efficiency per cycle and multicore are just superior solutions.
High performance cores are useful for problems that are hard to paralellize, but so far it seems that the breakthrough only occurs when a new approach to the problem makes it feasible on multiprocessing platforms (e.g. graph processing is hard to paralellize due to dependencies among graph nodes, Pregel and similar offer a different approach)...a 50GHz CPU won't save you if you need to process a huge graph (i.e. billions of nodes) on a single thread, it'll always take a lot of time.
As to the "record", I think IBM already had a Series Z that is over 5GHz.
Not hard limits, but yes, to my knowledge it is primarily physics that keep chips where they are. The requirements for power and heat dissipation start to balloon.
Ridiculous name. Maybe if they put a 'Go faster stripe' on the top of the chip people will believe it goes even faster!
Sounds nearly as fancy, makes the same amount of sense to the layman, and saves the rest of us a bit of time.
So a K5 PR-200 was actually a 133MHz chip, but it could match or exceed a Pentium 200MHz in some well-selected benchmarks.
At least the network switch makers have a sort of internal consistency, though you still have to learn individual vendors' styles.
AMD Exec 1: "Hmm, which superlatives can we use to sell this 'slightly more powerful chip'?"
AMD Exec 2: "Err... hmm... how about... ALL OF THEM?!"
AMD Exec 1: "Do you know, you're a freaking genius!!"
Still, it seems unlikely that no one had ever snickered about the phallic symbolism of the construction equipment of the term's original usage.
Those were the days, really, because there was still the possibility that in the future you'd have any damned thing in your computer, not necessarily an x86. You could have an Alpha or a SPARC or PPC or maybe an i960. And it would be silent and use no power and you'd install it in your bitchin' conversion van.
The Intel chips soundly win in anything else (encoding, Photoshop...), and by almost 2X in some of the single-threaded tests.
But at least you can say you got the bigger cache, clockspeed, core count, and debt than intel.
Most computers run on batteries these days, and those that don't drain ever more expensive electricity from the wall socket and at the same time waste a lot of it producing huge amounts of heat.
The more you get out of a watt the better. You can either trade in speed for lower power or trade in power for better performance, but in either case you want the performance/watt ratio to be the highest.
I would guess the power consumption of running the chip at 5GHz is pretty high. And running temperatures as well. And yet there are fewer and fewer of those huge tasks that you can only do with one core.
It depends on the workload, really. It should already be obvious that this part is not meant to be a Joe Everyman processor.
Hyperthreading keeps two threads "hot" in each physical core. When one thread is waiting on memory access, the core can do work on the other thread rather than sitting idle. (Memory access isn't that slow, so switches need to be fast to capture those otherwise-wasted cycles, which is why this is a CPU hardware feature rather than an OS-level software feature.)
Purely CPU-bound tasks [1] don't get any performance gains from HT. But almost all real-world applications spend a lot of time reading and writing memory, and memory access is pretty slow compared to CPU speeds, so in practice HT helps (otherwise Intel wouldn't have bothered to develop it and put it on their chips, which probably cost a lot of money).
> Have you check how disabling it influences compiling speed?
No. But I'd guess it would be substantially less than 100% speedup since they aren't actual, physical cores; but substantially more than 0% speedup since the compiler uses dozens or even low hundreds of megabytes of memory.
[1] By "CPU-bound" I mean register-to-register arithmetic. You might also be able to get away with hitting the L1 cache, which is a few KB, without triggering an HT context switch.
Doesn't have to be dimensioned that accurately. To a first approximation a 1% error in surface area would be about a 1% error in TDP.
I'd like to see very high temp CPU technology. That would be an interesting, challenging direction for hardware tech to move. A tiny lightweight 5 deg C/W heatsink is plenty if you're allowed to run at, say, vacuum tube redhot glow temperatures. I'm well aware of the solid state physics challenges of this, that's why I think it would be very interesting to see if anyone could pull it off.
Since the first half of the rumor came true, it's fairly likely the second half will too.