AMD reveals its first ARM processor: 8-core Opteron A1100
arstechnica.com
arstechnica.com
First I remember everyone driving the price up of an AMD video card just because bitcoin mining.
Second they got the backing in PS4 and Xbox One hardware.
Now an Arm 8-core CPU...although I find the clock speed (2GHz) kinda underwhelming, still AMD's pricing would entice me to buy 2 for the price of 1 Intel i7
Does not mean much. Their APU used in these consoles is very weak compared to what we have in mid to high end PC graphics cards. They were obviously chosen for their price there, not for the performance (which is MEH at best).
Also, I'd like to see your source for the statement that they are making money. I'm sure they aren't as huge loss leaders as they were in the past, but I'd be surprised to know if they are making money on them. If Microsoft could cut the MSRP, they would have.
If Microsoft could have, I think they would've offered their subsidized consoles (like they did the 360), and launch with Halo as their premiere title....
(this will probably get a lot of hate) then again Playstation actually decided to play video games this time and xbox one is trying to be a super Roku.
This was clearly mentioned by Sony before the launch of the console. As for the Xboxone, Microsoft also mentioned that if someone buys one game with the system, they are already making a profit overall. They are NOT selling their systems cheap compared to the previous generations, in terms of what's inside the box. I'll have to find the quotes, but I'm sure you can find them as well if you have 5 minutes.
Sorry but let me call BS on this one. We're not talking about PS3 kind of hardware here, where the architecture was unknown and new and developers had to learn a lot. The PS4 and Xboxone are using basically PC hardware under the hood, and the learning curve should be close to zero for most developers involved.
My point is that when the Xbox 360 came out, for example, the games running on it at launch were very impressive versus what the PCs could do at the time. In this generation, watching the launch games on PS4 and Xboxone just feels MEH at best. They are way too late in the game vs the power they are packing in.
In regards to the latest consoles, I have a ps4...and while most "anticipated," games are still on preorder, I don't think it's fair yet to call the hardware a joke, I don't feel the studios have taken full advantage of the hardware.....now the games offerings right now....yeah I agree, but I think the studios really compiled their games (at least the ones I play :: COD: Ghosts, and Battlefield 4) for previous gen consoles.
I am interested to see how the new halo plays though...just can't get past the fact that the xbox one requires kinect to be...ummm, connected to play.
Why? What's the point of doing that?
Though it is an interesting segue from this article about AMD supporting ARM chips (one of the biggest problems facing data centers is power density and with that heat removal) -- the PS4 has a total power consumption of 137W. That's the eight cores, GPU, 8GB of DDR5, blu-ray player...the entirety of the box consumes 137W.
The nvidia GTX770, undoubtedly a higher performance card, not only costs almost as much as a PS4, alone it consumes significantly more power than the PS4. That's excluding the rest of the box around it. That just isn't tenable for a living room gaming machine, which is exactly why consoles represent a necessary compromise. Similarly, several of the most interesting Steam boxes have GPUs significantly less powerful than the PS4's 7870, because you just can't drop a 450W device in the living room and call it a day.
I don't care if it costs more. It's available. Hardware better than the PS4 in consoles is NOT. That's the whole point. If I want to have better hardware than the PS4, I'll have to wait another 10 years for another console cycle to come. No way. In 2 years time the PS4 and XBoxone will be extremely weak even compared with low-end PC gaming rigs.
And you bet I'll want a high performing card (no matter what it costs) when the VR headsets come around (whether Occulus Rift or Valve).
Then these makers are missing the point. It will be more interesting to build your own Steam Machine then, and install SteamOS on it if no one is willing to put power in the living room. PC market is not console market.
Seriously did you miss that? http://www.forbes.com/sites/danielnyegriffiths/2013/11/19/re...
ANd that's just a third party evaluation, I have no doubt Sony gets way better hardware price deals when ordering millions of parts through their purchasing contracts, so of course they are making some kind of profit on each PS4 sold. They actually want to avoid what put them in the red with the PS3 sales.
> IHS iSuppli, having totted up the parts costs, has come up with a total of $372, with an estimated $9 labor cost bringing it to $381 – $18 below the recommended retail price.
So retail price is 399, out of which retailers make how much? 199 ? Tell me again how Sony is making a profit.
Remember that the clock speed is not necesarely everything. AMD may be able to get more work done with those 2Ghz than a Snapdragon 800 might using similar or less energy.
As far as I know, x86 (P6-based) is still king in IPC.
Then, a few years later I tried to install Gentoo on it. With X. And OpenOffice.
It chugged for days.
Learning exercise
>Why not compile on a different architecture thats faster or use distcc to make a compile farm? Last times I used Gentoo I just used the packages, took no longer to install than an rpm based distribution.
Well this was 10 years ago. I was 17 or so and I didn't know about distcc. I was young, dewy-eyed, and a huge noob.
It still feels really fast if you don't use X.
I lost that poor little machine when I moved across the country; not sure where it could have ended up. Now my only "system to screw around with" is my RasPi.
Cost of software and hardware for that was £0!
I still spend most of my time in a terminal window, but it's on a machine with ~1120x as much CPU horsepower and ~1638x as much RAM. Just calculating that made me go "wow" and realize how far we've come in the last 20 years or so.
All that said, I have a 16 core AMD server in colo that is running at about 3% usage across all CPUs, and yet it is slow as hell because the disk subsystem can't keep up (replacing the spinning disks with SSDs as we speak). The reality is that CPU is not the bottleneck in the vast majority of web service applications. Memory and disk subsystems are the bottlenecks in every system I manage.
So, I love the idea of a low power, low cost, CPU that still packs enough of a punch to work in virtualized environments. Dedicate one of each of these cores to your VMs; would be pretty nice, I think.
[1] http://ark.intel.com/products/75053/Intel-Xeon-Processor-E3-...
I agree that I/O and not the CPU is usually the bottleneck though.
Three generations back is still plenty fast for me. I'm running a Core 2 Duo on my desktop, as mentioned, and I don't see any reason to upgrade. I can't imagine needing vastly more in a web server.
I think it'd be interesting to see how these two product lines stack up on all the variables: cost, performance, power under load, etc.
works just fine.
http://ark.intel.com/products/77988/Intel-Atom-Processor-C27...
I have one of these sitting on my desk with an m-sata SSD. It compiles emacs just as fast as a R510 with 2 X5570s @ 2.93GHz.
When Intel goes 'tock' on these this Summer and manufactures them in 14nm rather than the current 22nm, the power consumption will drop even further.
and then there is Denverton. :-)
There's also going to be the Broadwell SoC, which will fit somewhere between Denverton and E3 v4's.
So it's technically possible that AMD could build ARM chips that are a lot more competitive (FLOPS-wise) with Intel than other manufacturers can.
What sort of issues with ARM memory/cache do you have in mind? These systems have been sufficiently powerful to keep the ALUs saturated on compute-heavy tasks on all recent ARM micro architectures with which I am familiar.
Of course it does in real life - unless you're working on very small amounts of data, cache level latencies (where Intel chips - non-atom at any rate - generally have much lower latencies) and cache pre-fetchers and branch prediction units (where Intel is generally 5/6 years ahead) can make or break the difference between the FP units being constantly busy or regularly stalled waiting for data.
There are more specialized HPC workloads (sparse matrix computation, for example) where gather and scatter operations are critical, and the efficiency of the memory hierarchy comes much more into play (but in these cases even current x86 designs are stalled waiting for data). There are also streaming workloads (which you seem to reference) where you have O(1) FPU op data element, which stress raw memory throughput and prefetch. However, one doesn't typically use these to make a general claim about which core is "more competitive (FLOPS-wise)”, precisely because they are so dependent on other parts of the system.
EDIT: looking at your comment history, you seem to be focused on VFX tasks, which tend to be entirely bound by memory hierarchy; even on x86 the FPU spends most of its time waiting for data. For a workload like that, you absolutely want to buy the beefiest cache/memory system you can, but that shouldn’t be confused with a processor being more competitive “flops-wise".
I have been having this argument since the 1980s in the school playground when kids with Spectrums would claim their 4Mhz Z80s were faster than 2Mhz 6502s...
With single-byte addressing, I'd say even a 1 MHz 6502 can make a 4 MHz Z-80 sweat. ;-)
And, for the MSX crowd, having to do 3 OUT's, one IN and an OR just to set a pixel is definitely ludicrous.
Commodores and Ataris of the time had very clever ideas about expansion. The intelligent peripherals are a brilliant idea and we should have done more of that.
I really hope MS ports windows server to it...
I think AMD are positioning this for the server market (though would love a nice cheap desktop with this chip. With that they are levridging there only real asset in branding and I feel it will not water down but help the Opteron range live on, given they are moving away from the x86 area.
I agree that re-using the name is bad mojo from a marketing perspective, too many people will be caught off guard by the lack of compatibility with the x86 chip set.
More importantly – will they provide zero-copy I/O like you can get with Intel network cards via their DPDK [1] or PF_RING/DNA [2]?
[1] http://dpdk.org/ [2] http://www.ntop.org/products/pf_ring/dna/
It doesn't matter. Modern processors have the complete opposite problem. It isn't that transistors are expensive, it's that they're so cheap you end up with too many and they generate too much heat. If you can stick a block on there which 60% of your customers can use and the other 40% can shut off to leave more headroom for frequency scaling, it's a win.
Also, the number of gates you need for a network controller is small.
As an entry into the server market, this sounds awesome; for things like storage aggregation and vm migration, gigabit ethernet is becoming a real bottleneck for many applications as the core counts has gone up.
Obviously I have no idea on the network controller or its sdk/driver support.
WRT to the opteron A1100 yes, I could see your comment. Something like a box of A1100 blades plugging 802.3ap to a common backplane, a trident chip there, and then a bunch of (Q)SFP+ northbound. A couple hundred gbs for around 150 watts of networking.
When I see A1100 I think an IO node with 10s of SATA disks attached. In that case Im only getting 10 or 20 per rack. A backplane makes less sense to me, running two DAC phys per box to a TOR I could see.
This is a microserver, designed to connect up I/O bound resources to each other. Imagine a cache like Squid running on this thing. Imagine mulitiple RAID-0 SSD drives on one side, and 20Gbps going out through the network.
This is NOT a computationally difficult task. For computationally difficult tasks, you have 8-core $2000 E5 Xeons (which get more and more efficient the bigger the workload you have).
However, filling your datacenter with $2000 Xeons so that they can spend 0.01% of their CPU power copying data from SSD drives to the network is a waste of money and energy.
The A1100 looks like it will be a solution in the growing microserver space. As Facebook and Google scales, they have learned that a large subset of their datacenters are I/O bound and that they're grossly overspending on CPU power.
Big CPUs -> Big TDPs -> higher energy costs.
This machine is designed with big I/O throughput (multiple 10GbE and 8 SATA ports on-chip), with the barest minimum CPU possible to save on energy costs.
The upcoming competitors to this market are HP Moonshot (Intel Atoms), AMD Opteron A1100, and... that's about it. Calexda's Boston Viridis has died, so that is one less competitor in this niche.
Either way, I'm hoping ARM64 will trickle up from iThingies to the desktop so I can buy a CPU with Virtual MMIO without paying an extra hundred bucks.
EDIT: And the whole Bitcoin-mining thing (or Litecoin/Dogecoin mining thing, these days), as mentioned in another thread.
err, can you link me to the other thread?
AMD doesn't, which is why you don't see much firepro, and why their entire line sells for so much right now while mining on hashes is big.
AMD cards are faster because they are built with more, simpler "cores", compared to Nvidia's fewer, more complex "cores". Mining benefits from the increased parallelization and doesn't benefit from Nvidia's "fancy" "cores". For sha256 mining, AMD cards gain an additional advantage by supporting a bit rotation instruction that Nvidia cards do not.
https://news.ycombinator.com/item?id=7141039
They just mentioned bitcoin mining and didn't go into any detail.
> Does AMD GPU hardware have a general advantage mining?
Yes. Disclaimer: I don't mine anything, am not that familiar with how Bitcoin works, and mostly heard about this stuff on the grapevine, so some of this is probably off. Corrections are welcome, if there are any cryptocurrency experts lurking.
At first, Bitcoin mining was done on CPUs. It uses SHA-256 for hashing, which (relatively speaking) isn't that difficult to compute. At some point, someone developed a GPU implementation for it, after which GPU mining quickly overtook CPU mining in terms of cost efficiency. The problem is easily parallelizable (just run a separate hashing procedure on each of the many processing elements in the average GPU), making it a better fit for GPUs than CPUs.
The short reason[1] that AMD GPUs are better than NVidia GPUs for this purpose is that while NVidia's GPUs use more powerful, but lesser in number processing elements, AMD GPUs use less powerful but more numerous processing elements (a processing element is basically marketing speak for each of the dozens or hundreds of small, specialized "CPUs" that make up a GPU). For mining Bitcoins, the extra capabilities of the NVidia processing elements basically go to waste, while AMD's cards, with more, less power-consuming elements, give you both "more hash for the dollar" and "more hash for the megawatt hour." This caused a huge spike in the value of ATI cards, completely unrelated to demand for PC gaming.
However, because this was a trivially parallelizable problem and there is big money at stake, miners came up with FPGA-based solutions (and later ASICs) dedicated to the purpose of mining Bitcoins, which in turn took the Bitcoin mining throne from GPUs. As I understand it, at this point mining Bitcoin with a GPU is a net-negative, and you need an ASIC farm to actually make anything off the operation.
At another point, Litecoin came around, and one of it's design goals was to be only practical to mine on CPUs, so that Bitcoin miners could make use of the underutilized CPU in their PC-based mining setups. It used the scrypt algorithm in place of SHA-256 for this purpose, which was designed to be computationally expensive and difficult to practically implement on FPGAs or ASICs (and therefore, more resistant against brute-force password hashing attacks), in particular by requiring a lot more memory than it would make sense to allocate a hashing unit on dedicated hardware. Unexpectedly, someone came up with a performant GPU implementation for that as well, giving AMD GPUs back the throne for mining (of Litecoins, and the Litecoin-derived Dogecoin). At this point, there is no sign of a cost-effective FPGA or ASIC-based scrypt mining device, so it looks like things will remain that way for a while.
[1] The gory details here, including something I wasn't aware of until now: NVidia GPUs lack an instruction for a certain operation necessary for SHA-2 hashing that costs them a couple of instructions each loop to emulate. AMD GPUs do have an instruction for it, so this automatically gave them another advantage over NVidia GPUs for mining purposes. Not sure if this applies to scrypt as well, but I'd guess so, since it was derived from SHA-2.
https://en.bitcoin.it/wiki/Why_a_GPU_mines_faster_than_a_CPU...
To make a real entrance into the server market, I would expect good virtualization support to be nearly a requirement.
Also there is an IOMMU implementation for supporting virtualization for IO. For example, the IOMMU and CPU MMU page table mappings are synchronized such that a DMA controller would also adhere to page table mappings set up for the CPU.
A recent post revealed some security problems using firewire (and a few other technologies) related to DMA[1]. Would the IOMMU features you're talking about prevent that problem?
You can have this protection, but then face programming issues if IOMMU and cpu MMU use different page tables. You have to update both. ARM IOMMU is designed so that it is automatically in sync with the CPU tables.
* Performance per dollar operating cost (performance per watt is closest to this)
* Performance per dollar capital expenditure (important for desktop systems, where operating costs are low)
* Performance per dollar TCO (sum of the above two)
The third one is the important one.
Why ARM?
How does the ISA impact the aforementioned criteria?
Why would a phones demand a different ISA?
ARM cores are typically slower in absolute terms than Intel cores, but at a given level of power, you can run more of them.
The differences between modern ARM cpus and modern x86 have less to do with the ISA itself and more to do with the way ARM cpus have been designed to be low-power for decades and have worked their way up the performance scale, while x86 has been designed for performance and has only lately been emphasizing low power. These lead to different design points.
The x86 ISA fundamentally takes more silicon to implement than ARM. More gates = more power.
This is not strictly true, the processor throughput also matters.
Total Power consumed = Power consumed by gates * Time taken to finish the job
Total energy consumed = Power * Time
For ARMv7 vs x86, yes, x86 just destroys ARMv7 (Cortex A15 etc.) in double (float64) performance.
While I do think x86 is still faster vs ARMv8, the gap is likely much less per GHz, because ARMv8 Neon now supports doubles much like SSE. Of course Haswell has wider AVX (256-bit) and ability to issue two 256-bit wide FMAs per cycle (16 float64 ops). Cortex A57 can handle just 1/4th of that, 4 FMA float64 ops per cycle.
That said, low to mid level servers are not really crunching much numbers. They're all about branchy code such as business logic, encoding / decoding, etc. Or waiting for I/O to complete.
So why would you care about math in a low end server CPU if it's not being used anyways?
This chip is interesting not because of the cpu core in it, but because it has two presumably fast 10GbE interfaces and possibility for a large amount of ram in a cheap-ish chip.
In fact not only regarding performance per watt, but also performance per dollar. It's just that ARM designs for lowest power consumption while Intel/AMD design for maximum perfomance.
Open-source software and the extreme efficiency goals of data centers make an interesting alternative to x86 now.
But overall ARM volume has been far higher than x86 volume for a long time even excluding all smartphones and tablets.
Most of our x86 servers at work have more ARM CPU's on them than they have x86 cores (most of the harddrives have controllers with ARM CPU's - some of them multi-core etc.). You'll also find it all over the place from washing machines to set-top boxes to microwaves. You find ARM cores in some sd-cards even.
I believe the projected number of cores for ARM last year was around 3 billion. I doubt x86 passed 500 million, which also means that both MIPS and PPC is competing with x86 for second place in number of cores for 32bit+ CPU's. (On the 16 bit or below end you also have surprises like 6502 derivatives shipping in ludicrous volumes)
So x86 has been "hot" for the market for main CPU's in devices consumers recognise as computers, and has been by far the most profitable architecture for a long time. Outside of that, though, it's at best at second place in total volume, and in most non-computer markets it's more likely to place in 3rd to 5th place in volume.
Linus has some interesting things to say about this too: http://yarchive.net/comp/linux/x86.html
Let's see: x86 code density is horrible for a CISC, there is hardly any advantage over ARM, which does great being a RISC. Also remember that the memory bandwidth is primarily a problem for data, but not code. ARM64 is a brand new ISA, it's the x86 ISA that is a relict from the times when processors were programmed with microcode. Intel is doing a great job to handle all this baggage, but to claim that the ISA gives Intel an advantage is ridiculous.
And finally, Linus has been an Intel fanboy since day one. Go read the USENET archives to find out. He received quite a bit of critique because the first versions of Linux were not portable but tied to i386.
> Also remember that the memory bandwidth is primarily a problem for data, but not code
RISCs, by design, need to bring the data into the processor for processing; but I see things like http://en.wikipedia.org/wiki/Computational_RAM being more widely used in the future, where the computation is brought to the data, and this becomes much easier to fit to a CISC like the x86 with its ability to operate on data in memory directly with a single instruction. Currently this is done with implicit reads/writes, but what I'm saying is that the hardware can then optimise these instructions however it likes.
The underlying principle is that breaking down complex operations into a series of simpler ones is easy, combining a series of simpler operations into a complex one, once hardware can handle doing the complex one faster, is much harder. x86 lagged behind in performance at the beginning because of a sequential microsequencer, but once Intel figured out how to parallelise that with the P6, they leapt ahead.
Linus being an Intel fanboy has nothing to do with whether x86 has an advantage or not. But even if you look at cross-CPU benchmarks like SPEC, x86 is consistently at the top of per-thread per-GHz performance, beating out the SPARCs and POWERs, and those are high performance, very expensive RISCs. I'd really like to see whether AMD's ARMs can do better than that.
Actually, Intel might be in a worse position with respect to vendor lock-in. I'm guessing a lot of early servers' lower layers like OS, webserver, etc. were proprietary; convincing the vendor to support x86 would have been a hard sell; and porting your application to an x86 environment was difficult.
All of these things would have had a tendency to lock people into their existing hosting choices.
Nowadays most servers run mostly / completely FOSS (at the lower layers) that can be easily ported to ARM. I'd imagine porting code to x86 from VAX or DEC or mainframe or whatever, was a lot more painful than porting PHP, Django or Ruby web apps to ARM today.
Of course, Intel does have deeper pockets and much of the desktop market, and may well be able to use that to keep ARM in check despite the fact that switching CPU architectures is probably much easier for website owners today than it was when Intel was trying to break into the server market.
See here:
http://m.bizjournals.com/albany/stories/2006/06/19/daily53.h...
http://finance.yahoo.com/news/-1-Billion-In-NYS-Funding-For-...
http://www.saratoga.com/news/amd-chip-plant.cfm
http://en.wikipedia.org/wiki/Luther_Forest_Technology_Campus
http://www.xbitlabs.com/news/other/display/20130904235756_Gl...
AMD has actually donated some nice amount of engineer-hours to enable ImageMagick to use OpenCL. And the difference is massive.
So if you have a webservice that allows users to upload images and then subsequently processes the images it's cost efficient to purchase few AMD APU's instead of massive Intel Xeon server.
There are a lot of other things to do, but Imagemagick is something that has support for it right now.
Perhaps I/O is the key here, N of these A1100 CPUs can easily saturate N x 2 x 10GbE, a single box with 64+ cores probably cannot push 16 x 10GbE.
Will server makers buy it? That remains to be seen.
Making a dev board available (let's hope it's also cheap enough to make hobbyists buy it) is rather clever. Without software tuned for it, the chip could fail on the market like Sun's Niagara and Intel's Itanium did.
Clock for clock, pipe for pipe (15 stages in A57, 14-19 in Haswell) x86 still wins because of per-instruction operand volume.
Though i wonder how well this chip would perform at 50w. Double the watts, maybe another 1.5ghz, might be viable for a cheap htpc.
As a general rule, power varies with the cube of frequency. 50W will only get about 2.5 GHz.
Does Qimonda ring a bell to anyone?
x86 is at risk here.
The fact it didn't leave AMD in charge of the x86-compatible market owes more to their superb Israeli chip design team being able to pull them out of the frying pan.
Within a 25W envelope??? I thought dual 10GbE chips consume 10-15W.
As long as I don't do computing intensive stuff like playing games or visiting some websites that overuse javascript the MBA runs really cool. Cooler than my body temperature.
-- typed on my MBA, lying on the couch, having it placed on my belly ;)
edit: found the link http://www.theregister.co.uk/2013/12/16/google_intel_arm_ana...
As much as I wish AMD getting ahead, this is not good news for efficient ARM servers.
I thought it was OK, getting the most performance vs keeping the TDP within passive cooling territory.
8 SATA-3 ports and two 10-GbE ports will run you a lot of power. I can't think of a single SoC that supports that much I/O.
Rounding out the SoCs, they'll also include dedicated engines for cryptography and compression.
"They" have ruined it for me...
As far as I know all crypto algorithms are deterministic based on the key/IV and the data.
Tampering with the RNG probably provides the best value for an attacker, and is harder to detect.
And if the crypto engine is compromised, how little/great of a leap is it to believe there is microcode to backdoor a general OS or crypto library?