AMD ships 16-core x86 3ghz CPU
maximumpc.com
maximumpc.com
I know it's a long shot, but could they try to go the way Intel did with Itanium, and invent a new architecture/instruction set specifically designed for server workloads? People have broadly expressed interest in ARM, but the support isn't there right now. If AMD had a new RISC architecture with awesome support and incentives for developers to target it, they might be able to steal share from x64 and ARM chips.
With desktop PCs continuing to lose popularity, and small form factors being popular with the desktops that are selling, that's going to be a much bigger problem for AMD going forward than their single-thread performance is.
If you can settle for "only" 2ghz, you can get a 16-core AMD 6366HE for $600 ($38 a core)
Bring on the price wars.
For the same money you can get a six core Intel E5-2630 (2.3/2.8GHz).
The benchmarks [1] say the Vishera (desktop Piledriver) has about 50% of the Intel performance per core at the same frequency, so unfortunately the 6-core Intel is roughly equivalent to the 16-core AMD.
[1] http://wccftech.com/amd-vishera-fx8350-x86-piledriver-pitted...
There are also power concerns with the AMD stuff, which matters a lot more in data centers/"cloud" environments than it necessarily does in a desktop.
Cache contention/invalidation is probably the bigger potential issue. I don't have a lot of experience with server VMs and I don't know how smart (or dumb) they are about pinning things to particular CPU cores in order to minimize cache issues.
I don't want AMD to go away either, but Intel's pricing won't necessarily go through the roof if AMD goes away.
The x86 CPU market is essentially saturated at this point. You can buy a "fast enough for most stuff" desktop computer for $50 or less at a thrift store, and most of us on Hacker News probably already own 4 or more x86 cores.
At this point, Intel's primary competitors are its own previous CPUs. I own several Core 2 Duos, a Core 2 Quad, and a Core i7 quad.
By all accounts, I am exactly the kind of customer they're going after with their newer CPUs. However, since I'm so happy with my current stuff, Intel is really going to have to outdo themselves (and price it right) before I'm moved to buy another Intel chip.
That's what I mean by competing with themselves.
Considering that power/cooling/etc in a datacenter often goes for north of $250 per Kw per month, you can bet that we watch performance/watt very closely.
Maybe a good question to go alongside this: if two teams of equal capability with equal access to fabs, patents, etc were to both start completely fresh, one team making the best x64 chip they could and one team making the best ARM chip they could, how much difference would there be in speed and power-consumption between the end products?
Although the talk is from 9 years ago, the material covered is still very relevant today. It is also quite funny. One of the things talked about is the processing of instructions and chaos theory, including non-intuitive stuff like inserting delays to make things run faster!
It should be noted that x86 processors haven't executed x86 instructions since the early nineties. They are translated into risc like micro-ops. The translation takes a very small fraction of the die area and it gets smaller which each generation of chip/process.
The difference between x86 processor implementations and ARM is that x86 try to get the greatest throughput possible. This is done via techniques like having multiple execution units and executing instructions in parallel where possible (known as ILP and typical values are 2.1), executing instructions out of order where it doesn't make a difference to the results, executing multiple instructions in stages concurrently (pipelines), having tracking for branches to better predict if they will be taken, speculative execution of both parts of a branch at the same time and throwing away the one that turns out not to be taken, complex memory machinery to keep code and data flowing, high clock speeds for the die as a whole, and even higher ones for parts if not all in use and the list goes on. This is not a requirement of x86 implementations but is what most of them do. Intel goes very far down this road, AMD not quite so far, and some implementations like Atom do barely any of it.
ARM processors generally do none of that. It keeps them smaller and simpler, which means lower performance and less power.
For your final paragraph, the instruction set is largely irrelevant. While x86 does have some warts, ARM does too (eg condition codes). The thing you left out is compilers as they generate the code to be executed. Roughly speaking the answer is the winner will be whoever has the better compilers. BTW Moore's law predicts transistor doubling per area every 18 months which most paraphrase as performance will double every 18 months. Someone did a study on compilers and found that compilers double performance about every 18 years!
Because ARM execution has been so simple for so long (eg no concurrent execution of instructions) the compilers haven't mattered that much. In a maximum performance world they matter a lot more, especially with instruction scheduling. And of course most programs would have been compiled a while ago, and probably use conservative optimisations (more aggressive ones can introduce bugs). With Itanium Intel had the idea of making the chip 6 way parallel and let the compiler figure out how to use that (ie smart compiler, dumb/simple chip). It didn't really work.
The tl;dr of it is that today, the RISC vs CISC or ARM vs x86 debate doesn't really matter - that's just a question of what "APIs" (if you will) are exposed from the CPU. Internally, they're all approaching working the same way.
The big difficulty I see is designing an instruction set, architecture and compiler which can show serious performance gains for workloads that matter. Making it easy for developers to come over from x64 while juicing the performance per clock would be a big win.
P.S. I'm not a CPU expert by any means. My understanding is that x86_64 has a big legacy overhead from implementing backwards compatibility with x86, and a goal of Itanium was to reduce the amount of the die dedicated to backwards compatibility. If anyone wants to correct me, feel free.
Probably not much right away, but over time, you'd see a real difference. One key issue is that x86 decode logic takes a substantial amount of engineering time to design in each generation. My spouse is a design engineer at a CPU maker and one thing I've learned is that there's a lot less automation than you'd think. Each new generation requires a bunch of dedicated engineering resources to make the extra-painful instruction decode work on a the new process with new performance constraints, etc. Decode is not something you can design once and then just reuse indefinitely; you take the NRE hit on every design cycle.
Spending time on that might be fine if x86 ISA was getting you a significant performance advantage, but since it is not, the extra NRE you blow on physical optimization of decode logic is just wasted effort that could be better spent elsewhere.
Anecdotally, I have an elderly neighbour who bought an i7 laptop, not because it was the higher number, but because her friend had told her that it was higher quality. Inadvertently, this tech-illiterate person had inferred that an i3 was somehow going to fail sooner or produce inferior results, because of the marketing Intel had performed. This kind of branding is far more powerful than Ghz nowadays, and it's a story that Intel more or less gets to make up. The only people technical enough to bother looking for a clock speed now probably understand all of the marketing jargon, more or less.
To return to the original point; Intel's marketing isn't dictated by what people want, Intel's marketing dictates what people want. They aren't trying to make higher clocked chips to convince people they're better. They have a dominant market position.
Obviously, technically sophisticated people understand that clock speed is not the sole determinant of performance, but a lot of people making purchasing or marketing decisions aren't that sophisticated.
I've probably gotten in past my depth at this point, but it seems like the bigger complaint is memory latency and bandwidth. My understanding was that this was part of the move to on-die memory controllers and increasing levels of L2 and L3 cache.
No, not necessarily. If one company puts out a lower clocked product with equivalent performance but lower power, the other company will be able to crush them in marketing and sales. No system integrator wants to try and sell a 1GHz product to the public. No one wants to convince retailer marketers that a system clocked at half the speed of their competitors is actually just as fast.
Take a look at laptop ads and ask yourself why they mention clock speed at all. That number isn't really comparable across different product lines or generations within the same product line. But people use it as a proxy for performance, so the ads keep including it.
The graveyards are full of companies that put out better technology products than their competitors.
> No system integrator wants to try and sell a 1GHz product to the public.
AMD has pushed lots of low power parts for mobile, for integrated systems, etc. If AMD had the resources to produce a low power, slower clocked chip, why was Turion such a dog? AMD hasn't actually done very well in the mobile space historically, when they could have produced low-power Ultrabook-like designs. Hell, they could've made something like the new Chromebook and had no fans. Either there's a severe lack of vision, or this is actually much more difficult to implement with the x86 instruction set than you're lettingon.
> But people use it as a proxy for performance, so the ads keep including it.
People generally don't care about clock speeds at this point, and I don't know if they ever really did. I worked retail about 6 years ago, and customers had no clue about clock speeds. Frankly, they were mostly worried about hard drives and screen size.
> No one wants to convince retailer marketers that a system clocked at half the speed of their competitors is actually just as fast.
Retail is the tip of the iceberg. HPC is a big market, commodity servers are a huge market. Halving your power consumption in those areas would be massive, and would give AMD a real cash injection. But there's no silver bullet there. You may be slightly right, but you're massively overstating the benefits compared to the costs of implementing it.
Because not even Intel can maintain two different architectures at once and stay competitive. AMD would surely go bankrupt before they could complete a major architecture re-design.
They don't have to. Reviewers would shout from the rooftops that your new laptop/tablet does not feel hot when holding it and lasts significantly longer on a battery charge. Then, you market your devices by quoting the reviewers.
Even on a desktop, a cooler CPU has advantages. Put it in a smaller, quieter box, and advertise that.
>they can never come down because that would decrease clock rates and lots of people have been trained into believing that clock speeds indicate performance
is a bold statement that is not likely to be true. Casual metrics of performance change, and I doubt there are many people confused as to whether they choose a 4GHz Pentium 4 over a lower clocked Core 2, much less something like a Xeon E5.
Then again, we've just switched to an entirely virtualized infrastructure running under KVM a few months ago, and ARM doesn't really have much in the way of hypervisor support. So, I guess there is some part of the software stack that ARM won't work for.
Fedora is available for ARM, which means RHEL and CentOS almost certainly aren't far behind. There's also a port of RHEL to ARM called Red Sleeve, which means most of the hard work of a port has already been done. Given that the potential cost savings are pretty big in a large data center, I suspect we'll start seeing deployments pretty soon, and the pressure on Red Hat to provide ARM will grow.
I'm not really arguing with you, per se. There are plenty of people who won't make the jump until Red Hat does. But, there are plenty of people who take their cues from other providers.
Anyway, hardware availability is an important prerequisite. Sure everything mobile uses ARM, so the Linux kernel is in good shape. But once there's cheap server hardware available, more work will be done on the distros.
Are you sure that you don't mix it with single-thread performance or per-core performance?
In addition to power concerns of many fast cores on one chip, server chips (the multicore champions) tend to be larger dies, and place higher value on MTBF. I believe both of these factors work against super high frequency server offerings.
I could be wrong on this, but I believe IBM's AIX line is not cheap, if you see where I'm going with this.
And that is imho a sign that it's not just a power consumption / clockspeed tradeoff, but instead there are actually limits on how high you can go at room temperature.
Of course you cannot compare 1ghz in 2003 to 1ghz in 2013, way more efficient per clock cycle with the right code.
Take Pentium 4 - which used the architecture Intel was working on when they were predicting that they'd have 10GHz CPUs by now. Passmark's score for a 3GHz Pentium 4 is 384.
An i7-3940XM runs at 3GHz, and gets a score of 10,490. That's a quad-core computer, so I'm going to go ahead and be sloppy by dividing by 4 to get a score of 2622.5 per core.
In other words, a single i7 core is doing almost seven times as much work per clock cycle. And assuming they only ramped up the clock rate, the P4 would have to be running at over 20GHz to match what the i7 is doing on a per-core basis.
Given the heat dissipation problems that come into play with high clock rates, that seems like a doubtful proposition. In hindsight, it looks like Intel definitely made the right choice by dumping the NetBurst plan and instead letting clock rates stagnate (they've even retracted a bit from the high water mark) while designing the processor core to do a lot more with each cycle.
Seems to me memory getting 10x more bandwidth and cache getting 8x larger probably accounts for the majority of the performance differences between these processors not instructions per cycle, which I think has gone up by more like 2x.
So a Pentium 4 with current memory and cache would need to be more like 8 Ghz if it scaled linearly like that (P4 had ~3 instructions per cycle, i7 ~7-8).
At 10Ghz, something moving at the speed of light can only go around 3 centimeters per clock cycle. Take out gate delays, and you wind up with something that just isn't practical.
The important things that changed were that smaller cores tend to leak more, which means that voltage had to be scaled down aggressively. And drive currents are now limited by velocity saturation.
The T2 processors from Sun had over 100 threads per chip. Its a shame that they didn't do better in the market place.
Our future CPUs should be tiny cubes, not flat chips.
Distance to each point in 3 dimensions is shorter.
Now turn them sideways, so they rest on the edges. Put the stack in a small ceramic container with copper bottom and top. Fill with a high efficiency thermal transfer fluid, and make sure the convection currents flow properly.
There you go, small matter of engineering.
And you probably still shouldn't. Quantum computing only provides a speedup for a very specific set of problems. For the vast majority of everyday computing tasks, quantum computing offers no practical advantage over classical computing.
This is what change looks like. Our intuition can be correct yet we still have to be open enough to recognize it when we it actually happens.
OTOH, quantum computers continue to not ship as a product.
On a side note: does anyone know what happened to the "Zurich" 32XX Opterons?