Google's POWER8 server motherboard
plus.google.com
plus.google.com
They've also formed a consortium to promote this processor, of which Google is a flagship member (http://openpowerfoundation.org/). The expectation (or hope, or fear, depending on your point of view) is that Google may be designing their future server infrastructure around this chip. This motherboard is some of the first concrete evidence of this.
The chip is exciting to a lot of people not just because it offer competition to Intel, but because it's the first potentially strong competitor to x86/x64 to appear in the server market for quite a while. By the specs, it's really quite a powerhouse: http://www.extremetech.com/computing/181102-ibm-power8-openp...
Well, the Power architecture had some success in Apple products, but ended with the inability of IBM to scale production and produce parts that consumed less power
Power8 has 230GB/s of bandwidth to ram compared to a Xeon's 85GB/s. That's nearly triple (270.5%) a XEON's I/O speed.
Further, a benchmark like this is complicated. How is RAM divided between sockets? What's the bandiwdth between a CPU and memory in another socket? etc etc
But I'll respond anyways. It looks like you have GB and Gb per second confused. One is 8x the other.
admin-magazine: claims 120Gb/s [1]
intel claims: 246Gb/s [2]
Independent claims on intel forums range from 120-175Gb/s [3]
Intel's cut sheet for their own latest generation xeon states it only supports 25GB/s (200Gb/s) memory bandwidth [4]
[1] http://www.admin-magazine.com/HPC/Articles/Finding-Memory-Bo...
[2] http://www.intel.com/content/www/us/en/benchmarks/server/xeo...
[3] https://software.intel.com/en-us/forums/topic/383121
[4] http://ark.intel.com/products/75465/Intel-Xeon-Processor-E3-...
The page you cite here: http://www.intel.com/content/www/us/en/benchmarks/server/xeo...
shows a triad bandwidth (STREAM) of 246,313.60 MB (megabytes) per second which is 240 GB (gigaBYTES) / sec.
If I'm making a mistake I'm sure we can work it out.
That is indeed impressive. no wonder google is considering these. memory bandwidth is indeed critical.
EDIT: Digging a little more gave me the answer: "The POWERn family of processors were developed in the late 1980s and are still in active development nearly 25 years later. In the beginning, they utilized the POWER instruction set architecture (ISA), but that evolved into PowerPC in later generations and then to Power Architecture. Today, only the naming scheme remains the same; modern POWER processors do not use the POWER ISA."
Gives a nice visual presentation of what happened.
I'm glad they are finally doing this, not so much because I care about what happens in the server world, but because so many product chip decisions at Google have been political (by choosing Intel chips) simply because Otellini was on their board. Hopefully this will signal a change from that.
I am going to guess this Dual CPU variant will be aiming at Intel Xeon E5 v2 Series. The 10 - 12 Core version cost from anywhere between $1200 - $2600. Although Google do get huge discount for buying directly from Intel and their volume.
Assuming the cost to made each 12 Core POWER8 to be $200, that is a potentially cost saving of $1000 per CPU, and $2000 per Server.
The last estimate were around 1 - 1.5 Million Servers at google in 2012 and 2M+ in 2013. May be they are approaching 3M in 2014/15. Even with most of those are low power CPU for storage or other needs. One million CPU made themselves could be savings of up to a billion.
Could this, kick start the server and Enterprise Industry to buy POWER8 CPU at much cheaper price? And Once there are enough momentum and software optimization ( JVM ) it could filter down to Web Hosting industry as well.
In the best case scenario, this means big trouble for Intel.
In the places this fits, it could offer substantial improvement. for example 10-100x performance/cost+power for in-memory cache servers.
And they're working on making this tech programmable while still keeping this same cost levels.
And all this in the context of moore's law grinding to a halt. So definetly ,intel will have a hard time ahead.
EDIT: it appears that the power8 support an open extension interface to other chips(CAPI). Which means will see such accelerators sooner than later.
The platforms team at Google is also developing an ARM solution. They have GPU's in testing as well.
http://www.extremetech.com/computing/181102-ibm-power8-openp...
I wonder how non-Google-scale developer could even potentially get to use POWER-based servers. Will they be available from the regular dedicated server hosting companies? What OS could they run? RHEL does support POWER platform, but for a hefty price: https://www.redhat.com/apps/store/server/ CentOS doesn't, presumably because all the POWER hardware CentOS developers could get is either very expensive or esoteric. That likely means I don't have to consider using POWER-based servers for at least 3 years, right?
http://www-03.ibm.com/press/us/en/pressrelease/43702.wss
and
http://www-03.ibm.com/systems/power/hardware/s812l-s822l/bro...
I can't stand it how their 'buy now' link for a product with a listed price then links to a 'get a quote' form. If they didn't do stuff like that I might have bought one of their machines instead of the HP that is currently churning away happily (32 cores, 192G of RAM, quite the little beast).
Which model and configuration of HP server did you buy and for how much?
I actually had popular dedicated server hosting providers in mind, e.g. Leaseweb. http://www.leaseweb.com/en/dedicated-servers
Re: linked servers. Thanks for the concrete info! $8K for 10-core / 3.4 GHz POWER8 with 32 GB RAM and 2x300 GB 10K rpm drives.
Those have to be some freakishly good 10 cores, to justify that kind of a price at least for _some_ use cases.
DL 385 G7 iirc, cpuinfo gives 2 times (AMD Opteron(TM) Processor 6274), max ram is 256G
Since POWER8 is little endian now it's pretty easy to get things running on it. We had Ubuntu ported in one cycle and 14.04 runs sweet on POWER8. All the compilers and the entire toolchain is ready to go. Everything in the archive works, just apt-get install.
All of your Linux workloads will probably just work on a POWER8 server. I started working on this server about a month ago, and have never used power-anything before. I just ssh'ed in did my work, and unless I did a uname or noticed the URLs with the arch in them when upgrading, it acts just like my Ubuntu x86 machines.
Yesterday at IBM Impact we deployed SugarCRM /w MariaDB and Memcached, a Websphere petstore, and Hadoop (using IBM's Java), all at once from zero to fully deployed and serving in _173 seconds_. These machines are _fast_.
Disclaimer: I work at Canonical and helped run the demo backstage during the POWER announcement.
http://cdimage.ubuntu.com/releases/trusty/release/
It's not mentioned on the website because the hardware is not publicly available yet.We have announced it on the blog though:
http://insights.ubuntu.com/2014/04/28/the-ubuntu-scale-out-a...
When the machines start shipping in real life (I think they said June?) it'll be more obvious on the main site.
All the surrounding ecosystem bits around Ubuntu will also get POWER8 support, so PPAs will start building POWER8 binaries, and all of the deployable services available on jujucharms.com will be available as well.
Also, Ubuntu works better on Power than CentOS.
Tyan boards, RH, SLES, Ubuntu.
> I wonder how
http://www-304.ibm.com/events/idr/idrevents/detail.action?me...
IDK when POWER8 get's integrated into Bluemix, but it will as systems get out there.
Virtual Loaner Program:
http://www-304.ibm.com/partnerworld/wps/servlet/ContentHandl...
On top of that, so much stuff just doesn't scale well. On the early Niagaras, even ssh-ing into the machine was noticeably slow. Oh? your crypto doesn't use all 64 threads? Hard luck!
An NVIDIA chipset and GPU would be able to go well beyond what NVIDIA is able to do with Intel chips (limited to PCI hooks).
For example their System/360 compatible mainframe CPUs are doing 5.5 GHz now. https://en.wikipedia.org/wiki/IBM_zEC12_(microprocessor)
Chip design is somewhat like software, adding more people or resources doesn't necessarily make your product superior. Look how long tiny AMD has been putting up a fight.
Many of Google's workloads are embarrassingly parallel. I used to work on Google's indexing system, and one of the binaries I worked with had a bit over a million threads (across many processes and machines) running at any moment, with most of those threads blocked on I/O. POWER8 has a fair number of hardware threads per core, which should help with the heavily multithreaded style of programming used in many Google projects.
Google's datacenters use enormous amounts of power. Several of their locations are former Aluminum smelting plants, because Aluminum smelting also uses enormous amounts of power so the power lines are already in place. My team luckily happened to sit next to one of the guys who designs Google's datacenters and we happened to overhear him say something over the phone about not being able to make sense about enough power to power a small town just disappearing from our usage. One of the guys on my team asked him when this power reduction occurred, and if it had happened before. We worked out that the huge swings in power usage happened when we shut down the prototype for the new indexing system. Most of Google's datacenters are largely empty space because they are limited by the power lines and cooling capacity.
In another instance, we discovered a mistake had been made in measuring the maximum power drawn by one of the generations of Google servers. Google had a program designed to max out the systems, and they plugged a server into a power meter wall-wart and ran this program for a while. The maximum power usage was under-estimated due to a combination of the office where the measurement was made being cooler than a datacenter (electrical resistance of most conductors has a positive temperature coefficient in the range of temperatures found in working servers), the machine not being allowed sufficient time to warm up, and/or the indexing system being more highly tuned than the program designed to maximize server utilization. (I like to think it was mostly the latter, but I suspect the first two were the main contributors.) The end result was that cooling was under-provisioned in one of the datacenters. During a heatwave in the area occupied by the datacenter used for most of the indexing process, the datacenter began to overheat, so one of the guys on my team was getting temperature updates every 10 to 15 minutes from a guy actually in the datacenter, and adjusting the number of processes running the indexing system up and down accordingly in order to match indexing speed to the cooling capacity. When you're really truly maxing out that many machines 24/7, some machines will break every day, so the indexing system (as most Google systems) are tolerant of processes just being killed either by software fault, hardware fault, or the cross-machine scheduler.
Through a combination of realizing smart phones were really going to take off and they were all ARM powered and realizing how important total machine performance-per-Watt is to some major server purchasers lead me to invest in ARM Holdings in mid 2009. (This does not constitute investment advice. The rest of the market has now realized what I realized in 2009 and I don't feel I have more insight than the market at this point.)
They've masked all the chips with something black. Are they hiding chips they are using, or is this something for thermal dissipation?
Second: while many CPU cores (with enough IO) is great for large Borg map reduce jobs, I am curious to see if Google will develop/use better software technology for running general purpose jobs more efficiently on many cores. Properly written Java and Haskell (which I think Google uses a bit in house) help, but the area seems ripe for improvement.
I know Google don't have a standard rack setup, but still, it would make seens to have all the expantion ports the end of the board... No?
Edit: Also, it's notable that both Xbox One and PS4 switched to use x64.
Note that PowerPC, as used in the Wii, Wii U and old Macs, is not exactly the same as POWER, as in this announcement. The POWER architecture is used by IBM AIX and AS/400 servers, and by the PS3 in its Cell variant.
For reference, the PowerPC 750 derivatives in the GameCube/Wii/Wii U are in the same family as the PowerPC G3 used in Macs around the turn of the century. It is also related to the CPU running Curiosity on Mars. So yeah, although Nintendo is a big customer of the Power architecture at the moment, they're not really breaking new ground.
I'm not entirely sure they succeeded in their goals, but with how well SPEs are used in PS3 games, I'm not sure it matters.
Contrast this with the software a PC runs -- mostly compiled to be optimized for a "generic" x86 CPU. In fact, it may have been compiled many years before the CPU was even designed. There is a lot more scope for runtime re-ordering to improve execution unit utilization.
If the whole world ran Gentoo, commodity CPUs probably would be in-order too.
AmigaOne's from A-Eon (http://www.a-eon.com/) and A-Cube (http://www.acube-systems.biz/index.php?page=hardware&pid=7) both running AmigaOS 4, and optionally Linux (at least for the ones from A-Eon, not sure about the ones from A-Cube).
There are some embedded machines, but not many now.
Apple G5s are still serviceable and supported by modern Linux distros, and are almost free...
For example I don't expect cloud providers to have a huge marked soon for non x86 architectures. Well there are JVM or other VM users which in theory could not care, as long as you don't need some native library.
In the past the battle with intel had to be played by providing an alternative implementation of the x86 instruction set for precisely the same reason: legacy.
The mobile market proved you can achieve good performance with ARM and especially better power per performance. I really can't wait to see some more fights in this arena.
Sure, it's doable, I have a few ppc at home, but I know first hand stories of small/medium companies just not wanting to risk that.
They just have higher costs at maintaining some dependencies, custom builds etc. You never know when you will get some new version of something, like jdk8. Then you have things like missing Go compiler (yes there is gccgo for ppc but still not everything works the same).
It's perfect for enthusiast, it's ok for companies with strong investment in IT infrastructure. I just wonder what is the best way to convince those small companies that there is no problem. Perhaps the tools/distros etc are starting to mature at the right point and this will soon no longer be a big practical problem.
Other languages with more OS agnostic libraries tend not to suffer that much from porting issues except from bit fiddling code. There is however the availability issue as you mentioned.
> nobody is working on it very seriously.
[0] https://groups.google.com/forum/#!topic/golang-nuts/sDV6ZfhG...
Porting an operating system is hard. Porting a compiler is hard. But most applications on top of those are fairly straightforward to port, if not completely trivial.
Anyone old enough to remember the early browsers - Mosaic and Netscape - will remember how every platform was supported, regardless of CPU (including MIPS, SPARC, Alpha etc.) and regardless of OS (IRIX, AIX, System whatever, VMS, Windows, etc.). Furthermore you didn't have to wait two years for the non-Windows version to be updated.
Only when Microsoft made NT Intel only did you get this arrangement where software was locked to a given architecture/OS. Thank goodness Linux came along.
ARM is definitely taking market share from Intel on the low end, but AMD is just about the only viable competition intel has right now in the part of the market where they make the bulk of their income and margin.
a dual socket board, 500W on CPUs, 600W with everything else.. the power supply would have to be something special, but the biggest challenge there would be getting the energy (ala heat) back out of the box..
GPUs have similar TDPs and issues - that's why the HSFs on top of them are so massive (and hence GPUs have a bit of an advantage here - they have the entire PCIE board to fit their cooling hardware on)
finally, 4.5ghz? what the hell? in one clock cycle, a beam of light wouldn't even get half way across the board (EDIT: not chip). branch/cache/TLB misses may literally kill any reasonable performance you might hope to get out of it. intel get around this by having years of market leading research in branch predictors, caching models, etc. and it's going to be no mean feat to match that.
i know IBM aren't exactly new to this game. but AFAIK x86 has always been faster, clock for clock, than POWER.
that said, i hope my concerns are misplaced. i'm hoping intel get some competition in the server room. it will be of benefit to everyone.
I don't think 4.5 GHz is somehow ridiculous when 3 GHz is routine (and POWER7 was 4.2 GHz). Hundreds of cycles of latency when accessing anything off the chip is now routine - that's the world we live in now. I think that the biggest problem is that IBM is not able to make the investments (especially in semiconductor manufacturing) to match Intel's rate of bringing technology to market. The current POWER7 is a 45-nm device if I remember correctly, and this 22-nm POWER8 is not yet on the market. Intel has been selling 22-nm Haswells for how long now? And of course the POWER7 chips have been up against next-generation semiconductors for most of their life.
EDIT: I see that IBM started selling POWER8 systems a few days ago. That's close to a year later than Haswell, and what's more, this chip is likely to compete against 14-nm processors for most of its lifetime.
Light would travel about 660 millimetres in 0.22 nanoseconds
Important note: the wave propagation speed in copper can be as bad as .42cYeah, starting with the POWER7. POWER8 has 96 MiB of eDRAM (e for embedded).
The Centaur memory controllers also have 16 MiB of eDRAM, max them out at 8 and you get 128 total at L4.
Compared to Intel's current offerings, the L1 data cache and L2 unified cache are twice as big. Don't know about timings, though.
The biggest Intel Ivy Bridge Xeon server CPUs have slightly more transistors (100 million), but on a much smaller die, 31% less area. Look at the ones with 12 and 15 native cores: https://en.wikipedia.org/wiki/Ivy_Bridge_(microarchitecture)... they list at $2336 to $6841.
Hindsight makes Itanium look like even more of a disaster, when that energy in that era should have gone into evolving the x86 platform for the future. Without AMD doing what they did (x86-64) I wonder where Intel would actually stand in the server market today.
Keep the x86 decoder front end that they have now. Add another one for the better ISA. Add another mode that kicks the CPU into that ISA. Current x86-64 OSes already generally support two architectures: x86-64 and i386. This would just be a third one. Then everything could move to the new ISA incrementally, and legacy software could keep on working forever using the legacy decoder.
I'd guess that the x86 ISA is no longer enough of a bottleneck to justify it. Throw enough transistors at the problem and perhaps it doesn't matter anymore whether your ISA makes any sense.
Not only is it possible, but it has been done. Look up the Transmeta Efficeon. Note that it's not RISC-like but VLIW (which I like to think of as "the next thing after RISC").
If Intel Cpus use a Risc like architecture nothing would prevent static translation to it, which was not feasible for Transmeta cpus. The x86 software layer would then only be there to preserve backwards compatibility.
IBM was pretty close to cutting off most funding to future development, and they just closed the factory in Minnesota that made the servers. They are pretty much the last survivor of the Unix wars.
There are major customers using this stuff and scaling up, but the industry as a whole is shifting to scale out. There might be a good story here with POWER8, but can you trust that the platform will be around?
....which they moved to Guadalajara, Mexico. Given the money they put into Watson, I would expect they keep the POWER machines.
It's not something they spread around, but from what I recall from a Q+A with IBM engineers the Watson prototype was developed on AMD xSeries servers, and only moved to pSeries machines late in the development before the Jeopardy! matches.
The business growth for power was displacing PA-RISC, SPARC and Itanium... IBM is the last man standing, but are competing against whomever is supplying AMazon, etc for barebones x87. In the IBM line, it's jammed between commodity x86 and high margin zSeries.
Sounds kind of surprising even if IBM did some of the bringup work ahead of time, but maybe they've got little endian assumptions baked in many internal protocols/apps.
This isn't a concern for low level developers, such as the kernel developers - they understand the concerns and take care to implement code in portable ways.
The issue is with user-space developers who think C and C++ are a good choice of language, and they have no qualms using bitfields, unguarded compiler pragmas, violating strict aliasing rule, and failing to specify the endianness their protocols use in the protocol itself (BoMs are not universally used) - also there is often a failure to provide the endianness conversions in implementations of such protocols where necessary. Not to mention a complete lack of standard way to test the endianness of the current machine, which typically requires violating the strict aliasing rule to check.
Since most Linux distros fully support a very diverse set of machines, endianness is usually not a problem with most of the software that's already part of a Linux distro.
As for software developed inside Google, they hire smart people. They'll manage.
The big news here is official support for KVM on POWER. Use all your existing automation, Openstack, etc, unchanged.
Re endianness, Linux is comfortably bi-endian and so is approximately all portable Linux software. Certainly there isn't any "tremendous" difficulty, which is why I was expressing surprise at ditching the standard big-endian ppc64 userspace.
When people first started floating around the little endian PPC patches in 2010 it seems the motivation was some GPU hardware (http://lwn.net/Articles/408848/) but that doesn't really make sense in this case.
userspace, the root of all evil
Most architectures that support unaligned access have a small penalty for unaligned access anyway, and some architectures (Does anyone remember Netscape Navigator on Solaris SPARC crashing with SIGBUS much more often than the same Navigator release crashing on x86? At least Solaris 6/7 didn't include kernel code to emulate support for aligned memory access on SPARC.) don't support it, so it's best to avoid unaligned memory access in C code.
I don't recall the JVM specification forcing a particular object layout on an implementation, and I believe most JVM implementations naturally align all object fields rather than packing them for minimum space usage. I believe an implementation could reorder the fields in order to optimally pack them while avoiding unaligned accesses, at the cost of breaking any hand optimization of locality of reference made by the programmer. However, I think the space savings for almost all programs would be very meager.
What are some use cases for a server like this for Google? I'd love to see these available in the IBM Cloud (SoftLayer) but I think they will be too pricey and reserved for enterprise.
When I see stuff like this it is painfully clear that from a technological perspective a company like duck-duck-go has a huge amount of defensible moat to cross before they can begin to be a serious contender. Think about it for a second: the company that you're trying to compete with is operating at such economies of scale that it can afford to have its own custom motherboards + non-standard expansion boards made.