Power 9 May Dent X86 Servers: Alibaba, Google, Tencent Test IBM Systems
eetimes.com
eetimes.com
Lots of people in tech talk about being interested in trying it, it has a number of interesting characteristics, but the barrier for entry is fairly high.
It's possible to get a POWER8 based server on Softlayer, IBM's cloud product, but not as a VM. You can get one as a bare metal server, but you can't get one for an hourly fee. You have to pay for a full month, which starts out at around $1000.
There are very few individuals that would be willing to make such a commitment, but so many that would be willing to spend a few tens of dollars on spinning up a VM for a few hours to see if it provides value.
If you want people to get excited about it, or interested in using it, you really need to make it easy for people to test it on a small scale.
There are no entry-level POWER machines, SPARC machines. There is no entry-level IBM i or Z. There isn't even a software emulator for i (and even for Z, where there is a very nice emulator, you can't legally run anything remotely current on non-IBM hardware).
You can get a cheap POWER or SPARC on Ebay, with a sharp price drop when the equipment is EOL'ed, but that's no way to entice people to build new stuff on a platform.
I love my PA-RISC, my RS/6000's and my SPARCstations, but, neat as they are, they aren't anything that would make me consider building for their modern descendants. And I won't pay more than a decent, proven, certainly useful Xeon workstation worth of money for a cool, but exotic-that-may-not-run-my-software-well, POWER9 performance equivalent.
If IBM wants me to pay Xeon Platinum prices for their production gear, they'd better allow me to pay Core i prices for development gear, or else I'll just keep deploying on Xeon, which is good enough.
Typed on a very comfortable and reasonably priced Core i7 laptop.
I know of companies that regularly throw away second generation core i5 dual and quad core systems with 8GB of RAM.
The same software you develop on your $20 headless desktop PC can be easily moved over onto a $15/month KVM virtual machine somewhere in hosting. Or onto a used $250 1U server that you can pay to colocate somewhere for $65-100/month with bandwidth. All x86-64 platform.
The massive, massive economies of scale for x86-64 platform stuff are going to be very hard to get people to move away from unless there is an amazingly compelling reason. The ability to cannibalize random recycled computers to make one working computer out of "free" parts is a big thing for the developing world.
IBM has to compete with free hardware.
If you come from a world where lock-in is normal then you may well build something that locks everybody out.
Let's say that a 44RU rack of 1U, narrow (half of 17.5" width), dual motherboard, single socket systems costs $N if it's built with Xeon or EPYC CPUs. As an example here the Supermicro barebones which are two long, narrow motherboards in 1RU with single sockets and front to rear wind tunnel airflow. Same rack of systems costs $N multiplied by 2.25 if built with power9, but only performs 1.25 times faster.
I would very much like to be proven wrong so if you have access to some specific "this hardware spec costs this much money and you can buy it now" info, that i can compare to barebones supermicros, please do share it.
But I will say that your comment about $/MIPS isn't accurate. There are many people, including Oakridge national labs, who are using it because of the higher performance and v100 integration. Intel is obviously not too keen on letting Nvidia succeed.
I sit on the POWER Customer Advisory Board, what you have raised has certainly been something I've told IBM multiple times over the last few meetings I've had with them.
I plan to continue pushing on this front.
Even when they try some new trend ( aka cloud computing for the masses ) the economics aren't that good because the way they are structured, it's a culture thing.
It's like Intel developing ARM CPUs, they have the means, but they can't do it successfully.
Bootstrapping an alt arch is hard, it really needs a loss leader OR a significant performance delta (which they now have with accelerator workloads). Google deploying it in prod is a major milestone and will help with volume and confidence.
Supermicro has a P8 and P9 line. It is price competitive with x86. If you are interested, send their sales a note.
POWER is nothing like this. There are no development boards at all, for anything less than $thousands, and the real servers have great performance but sky-high prices.
You can web buy an S821LC straight from IBM for $5k for the past couple years. The AC922 is GA. Supermicro will sell you a P8 right now and P9 in May for nominal prices and the CPUs are cheaper than Skylake by a very wide margin.
Not directly related to your comment but I get the feeling a lot of people complaining about price are navel gazing and have no idea how much a production server costs and how the costs break down. Right now storage is generally 50+% of the cost. DRAM is a very high fixed cost at the moment as well. Intel flatted the quad socket SKUs into the "Scalable Series" so Skylake represents a big price increase for a lot of builders. To help re-calibrate people, IBM doesn't have $6-12k CPUs in the dual socket config but Intel does at bins people would want for common workloads.
They're not asking for production servers.
$5k for a workstation, for a minority architecture, is basically a fancy way of saying "no" to the army of tinkerers you need to widen the base.
"At STH we are working with Gigabyte and Cavium and will share more about the ThunderX2 architecture as we are given the go-ahead. We have heard the next production run is in the Q2 2018 timeframe."
This has been my experience with cavium, bait and switch slideware and when you do finally get a sample it has so far been undesirable (octeon, TX1)
Remember, it's never going to be the same price as a RPi.
https://www.pine64.org/?product=clusterboard-with-7-module-s...
The point of clusters like this is not the speed, but the fact it's a cluster, with all the bottlenecks a cluster has.
On USB it seems I'll have to route all traffic between the nodes via the cluster controller, which will also serve the shared NFS volume.
It’s more productive to think about how to harness lots of cheap ARM cores if intel isn’t doing it for you.
My employer has a few. It’s fast and looks impressive. But when you cut away the bullshit, it mostly exists because the sales guy presented a story where the cost of a new POWER box is a better deal than maintenance on the old one.
If you really dig into it, there’s no scenario other than a license play where a transition to Intel isn’t more cost effective. Even in those scenarios, you can usually engineer a solution (Oracle, etc) where you deliver a better ROI on commodity hardware or cloud hardware.
But yes... Vintage computing doesn't count.
I will concede one point: it was not easy to purchase. I prefer to host my personal site on Power hardware; I started with AIX in the 3.2.5 days and I ran Floodgap on an Apple Network Server 500 for the better part of 14 years. I wanted to get a POWER7 to replace it in 2010, I budgeted $15k for it, and IBM wouldn't take my money. I couldn't find _any_ IBM VAR who would do an end-user sale because I wasn't going to buy the service contract.
Eventually I found a reseller who was more than happy to take $10K of my budget for a decent 2-year-old POWER6. It had a backplane burp a couple years ago but otherwise has been pretty damn spiffy.
I'll concede IBM has to do a lot more to get these systems into people's hands to achieve a critical mass and it certainly wouldn't hurt to make them cheaper up to a point, but the systems are out there, and you can get them (and find them).
As a postscript, in my current job (a large local government agency) I was in the CIO's office one day and the regional IBM salesdroid dropped by. Just to needle him I told him this story and he gave me his card and told him to call him with any parts requests, any time. I still buy from the reseller, though. They've earned my personal business.
The other thing with power is that it has been a scale up vs. scale out product. That makes sense when you want to optimize your Oracle/SAP licensing or something similar. It doesn't make sense for modern use cases where you're using open-source or other solutions where growing infrastructure into your use case makes more economic sense.
”The Power Cloud that enables developers offers no-charge remote access to IBM hardware, including IBM POWER8, IBM POWER7+ and IBM POWER7 processor-based servers on the Linux, IBM AIX and IBM i operating systems.”
I don't think IBM cares about people using it, especially people who can't afford $1000 dollars. They are only interested in high margin corp/gov pork barrel deals and need something that they can plausibly claim is superior in order to charge 10X for it.
Plus IBM licenses the design these days (through OpenPOWER), so several of these players are building their own chips, boards, etc. Someone could inevitably enter the low-end market but lower-end devices have thinner margins and a lot more competitors.
Even right now you can get real chips to go on a board on pre-order, just under $400 (TALOS II preorders) -- the mobo is the pricier part, but part of that is likely due to the BOM choices on that piece from Raptor Engineering. A smaller form factor motherboard (maybe with 1 socket) could land in the sub $2000 range for a whole mobo+cpu -- which is certainly competitive with similar HEDT/workstation prices[1]...
[1] I just dropped $1,500 on a 1950X threadripper and associated mobo earlier this year, so this price range is certainly alive and kicking, IMHO.
Power9 has 120MB of L3 cache that's shared between all cores. Which means you can ACTUALLY have a full 120MB-sized problem set and share all that information between cores.
In short: Threadripper is great for sure, but its not necessarily a fair comparison. In an apples-to-apples comparison (ie: Monero Mining), it seems like a Power9 server is 3x better than Threadripper (Power9 gets ~3000 hash/sec, while Threadripper is roughly 1000 hash/sec).
https://www.phoronix.com/scan.php?page=news_item&px=POWER9-C...
> Using the xmr-stak-power PPC64LE-focused Monero miner, they are seeing great performance with it running on dual pre-production 16-core POWER9 processors. There's a hash rate of 2945H/s while this POWER9 system is pulling 350 Watts DC power.
And mind you: Monero / Cryptonight only uses 2MB of L3 per core. So that's practically Threadripper's ideal problem. Imagine if you actually had a dataset that was larger than the 8MB per CCX that Threadripper is limited to.
Not to hate on Threadripper at all. Its cheap and high performance. I'm seriously considering a Threadripper system myself. But these Power9 specs are incredible, and I'd definitely like to test one if I could afford one.
https://www.youtube.com/watch?time_continue=226&v=CS7M392Ia_...
I'm sure there are workloads that really do benefit from the huge unified L3, but it seems like a pretty small niche.
Each Power9 is consisting of either 2-super slices or 4-super slices, depending on which Power9 you get. (Corresponding to 4x SMT or 8x SMT respectively).
Threadripper has 4x integer pipelines, 4x floating point pipelines, and 2x load/store units (called AGUs by AMD) and supports 2-threads (aka: 2x SMT).
So Threadripper definitely is "broad", but Power9 is "broader". The 8x SMT Power9 can perform 8x loads / stores per cycle per core, while AMD's Threadripper can only perform 2x loads/stores per cycle per core.
IBM has extremely flexible hardware which, using IBM's hypervisor, can be configured on the fly to present itself in a number of different ways. For example: few powerful compute units, many medium compute units, or very many less capable compute units.
So right up front the idea (which a lot of people might assume without really thinking about) that a given chip has a fixed number of "cores" doesn't really hold true. More to the point the next question that presents itself "what makes up for a core and what determines how many a given processor has" is best answered with "it depends" (at least as far as IBM Power goes).
One of the main reasons this matters is that all sorts of businesses use the idea of "cores" as a fixed entity as part of their pricing structures... and now it's all muddied.
In regards to cores, the distinction is between the PowerVM and Linux (OpenPOWER/PowerNV) ecosystem variants. Both are made to process 96 threads, but the difference is in whether the thread processing units are grouped eight to a core (SMT8) or four to a core (SMT4).
The PowerVM version, some say for licensing reasons, gets the SMT8 cores. Either way, you get 96 threads:
12 cores * SMT8 = 96 threads 24 cores * SMT4 = 96 threads
I guess this is pretty much all just leverage to hold into Intel's face to get better prices out of them. "Look, we could totally convert to POWER..."
Glad to see them throwing in the towel on a pricey and proprietary component. Competition is good.
Sure, I could have bought one of the AmigaOne systems, which would have been somewhat cheaper, but those are basically embedded systems in ATX cases. The 12-year-old Quad G5 I'm typing this on would mop the floor with one of those. This isn't a slam on the Amiga community, who tried hard to bring in systems at not-completely-eyewatering prices, but this ends up producing boutique systems running underwhelming chips in unimaginative designs because of the small numbers and the rigid price points. More to the point, I don't have a strong Amiga history, so the AmigaOS side of those machines would be mostly wasted on me.
If there's going to be an architectural alternative in the grunt ballpark with x86, and I don't think ARM has gotten there yet, right now it's going to be Power ISA. And everyone in this thread complains about the cost. Understandable, but if no one steps up and buys one, no one will make any more of them and they will never achieve the necessary economies of scale. Fortunately I'm a PPC bigot^Wzealot with more money than sense. So I'll take the plunge. :)
The fact the firmware is auditable, no management engine/PSP crap, I get full schematics, etc., is just a bonus in my view.
It seems I have close to no experience with POWER (other than older Macs), so, I am a bit concerned about what I will and won't be able to do on it.
When I receive mine I hope to be able to use it for most things, but as a fall back it should be able to be used for some self hosting.
Any particular place POWER guys hang out? I haven't been able to find a community that talks about day to day off x86 (albeit I haven't done much looking).
Disclosure: I work on IBM cloud, but only tangentially with the Power systems.
Last I looked, Power cloud offerings required the customer to be a big company that buys a lot of compute.
As for the future, I can't say much publicly (le sigh) but let me just say that personally I think future of power is bright, especially with cloud providers.
e.g.: "Power 9 should do better than its predecessors given its costs, bandwidth, and ease of porting. Power 9 is IBM’s first to use standard DIMMs, opening a door to other standard components that are, overall, cutting system costs by 20% to 50% compared to the Power 8, said IBM’s partners."
The burden of proof is on those claiming the pricing does not present a barrier to entry to the ecosystem.
I am curious about developing for Power. Please link me to an entry-level system I can buy on the web.
Setting that aside though, its true that easy experimental access can allow a technology to "sneak in" to a market that previously ignored it, but that isn't really the case here. (nor was it the case for SPARC or Itanium) These systems are being built for folks who are going to deploy a lot of them and they already know their cost of ownership numbers for existing x64 boxes.
The next wave is probably running on commodity x86 hardware running Linux and using gamer-grade GPUs. They'll probably be first deployed to production on x86 virtual hardware in a cloud provider.
If you want companies to adopt your technology, make their developers adopt it first. That's how most of the current enterprise DevOps tools happened.
At least at Google the driving factor is total cost of ownership, full stop. When you need as much compute as they do to deliver on what they deliver, saving a few percent on the TCO flows right to the bottom line.
That said I don't see them going full on with this technology if its QP$S[1] isn't competitive with x64.
[1] Queries per dollar-second.
Where it can make a substantial dent in the x86 server is precisely this HPC niche. A 10% faster $300 box is worth $30 more, hardly enough to warrant extra work, but a 10% faster $10 million dollar machine is worth a million more.
Sun/SPARC for example, benefited greatly in Dot-Com v1.0 because so many universities had great prices via edu discounts. The sysadmins and other Unix users all cut their teeth on Sun boxes; so when they got jobs they took along their familiarity with Sun, which resulted in a lot of sales.
A. Many of the people who today specify datacenter hardware came into that career as enthusiasts (aka "early adopters") and are more comfortable recommending what they have experience with.
B. Enthusiasts bring the consumer volume that drives down production cost. This is a necessary driver of Moore's law.
If, however, IBM can get me a 32GB single-core/8-thread POWER9 machine for the same price Lenovo can give me a performance-equivalent 4-core/4-tread Xeon, I'll be tempted. Worst case scenario, I'll still run my x86 workloads on my trusty Lenovo and use the POWER9 as a nice X terminal.
x64 by AMD pushed Intel to do better. Things became faster / cheaper. AMD fell behind and Intel rested and priced for high end went up. AMD released Ryzen and all of the sudden we see movement from Intel again. ARM on the low end pushed Intel to look at their lower powered chips and try to do better. IBM pushing P9 to win at the top at a reason price point pushed on AMD and Intel. Competition is only good here and brings value to all of us.
All of this is only good.
The architecture was something I was very excited about for low latency workloads; it seemed that only thing stopping IBM from overtaking x86 in this space was the actual availability of hardware. POWER9 has been "in the works" for, what, two years now [1]? Google have been running POWER9 internally for at least 12 months [2], and IBM seem to have been caught up in the AI hype train by pivoting the POWER9 to some kind of "AI processing system" [3].
Had the P9 been released before Skylake, it might have been successful. By the time it is GA, it will be competing with x86 chips two generations higher than were available at the original P9 announcements and SPECInt benchmarks.
This has been made even more frustrating by the fact that Oracle effectively killed off SPARC during this time. Perhaps Fujitsu will continue to run with it, but I think it's fair to say that development will stagnate and support will dwindle as it is increasingly obscure and niche architecture.
[1] https://www.nextplatform.com/2016/08/24/big-blue-aims-sky-po...
[2] https://cloudplatform.googleblog.com/2016/10/introducing-Zai...
[3] https://www.forbes.com/sites/tiriasresearch/2017/12/08/why-i...
Google, on the other hand would need a reason to switch. Somewhere there's a spreadsheet with a number for energy cost savings that would make it worth it for Google to switch architectures. I've heard it's as low as 10%.
China has Zhaoxin, which develops domestic x86-compatible cores that have received massive state funding, despite little commercial success. They way it looks, I think that China has decided to stick with x86, with Zhaoxin as an "escape valve" for the case where chip supply from Intel and AMD gets threatened.
I thought they had a thing for MIPS64, at least in their supercomputers.
Google is already at least experimenting with POWER9, but it's not clear just how far that goes.
Here is the Google Blog post introducing Zaius: https://cloudplatform.googleblog.com/2016/10/introducing-Zai...
of all places you think Google does this kind of cost-analysis in a spreadsheet?
If POWER unlocks something fundamentally new and historically impossible on X86, like ARM unlocked cell phones, and such, maybe that’s another reason to try it.
At the end of the day, nobody got fired for choosing x86. Its well understood, it’s cheap, development tools are accessible, and it works.
https://wiki.raptorcs.com/wiki/Speculative_Execution_Vulnera...
You have to apply patches to the LPAR/OS, VIOS and firmware. IBM has acknowledged that there will be a performance hit but have not provided any quantified numbers that I have seen so far.
"All POWER8 and POWER9 results in this table reflect performance with firmware and Operating System updates to mitigate Common Vulnerabilities and Exposures issue numbers ... known as Spectre and Meltdown"
I also just found the rPerf and CPW consolidated spreadsheet from IBM (including pre and post Spectre/Meltdown numbers).
https://www.ibm.com/developerworks/community/files/basic/ano...
For those like me whose comprehension was hindered by all the upper casing.
Not a chance.
MtG didn't come out until 1993. So yes, it's a co-incidence.
Where Power8/9 really shines is memory bandwidth - it's orders of magnitude faster than leading edge Intel stuff currently. But if you're not memory constrained, Intel machines can still have a bit of an edge.
That makes Power sound very appealing. I've never met a real life scientific simulation that wasn't often constrained by memory bandwidth. I'm sure they exist, but they aren't all that plentiful. I'd love to try running the code that matters to me on Power. (For one thing, I don't even know if it would pass basic regression tests; numerical code is touchy. Are gcc/gfortran reliably adequate on Power like they are on x64?) Unfortunately, as everyone else is commenting, entry level x64 is dirt cheap and entry level Power isn't.
I'm more hesitant to say the same about clang/llvm but it's been awhile and I'm probably not current.
I'm surprised if numerical issues are that much of a worry. x86 has now caught up in doing FMA, in particular. I don't know what the build systems actually are, but most of Fedora (and Debian?) is available for POWER, and will typically have at least minimal tests in the package. It's not as if POWER is new for HPC; the Daresbury HPC-X machine was quite high in the top500 15 years or so ago, as an example I was physically close to.
I'm sure there are a lot of problems that fit inside of 120MB.
"Processors have two modes of thermal protection, throttling and automatic shutdown. When a core exceeds the set throttle temperature, it will start to reduce power to bring the temperature back below that point."
https://www.intel.com/content/www/us/en/support/articles/000...
If ARM or Power were faster than x86 we would be using it right now on servers, most of the top500 HPC clusters run on x86_64 ( Intel Xeon ).