IBM Opens Power Microprocessor Architecture
bits.blogs.nytimes.com
bits.blogs.nytimes.com
It means IBM doesn't want to do HARD engineering anymore. They just want to do sales and handling of stuff for high markup. This is the last really cool thing to go before the finish dismantling that former technical giant.
POWER was never about flops/watt it was about the IBM ecosystem. It's calling in for support on your hardware/software and getting elevated to somebody in the same department that will actually write the fix. It's having actual field engineers that went to the factory and learned how the stuff is put together. It was never cheap... But it was good.
ibm's POWER hardware is the last of the "old iron" lines. X86 stuff just doesn't hold a candle in the "well built" category. My company went from an "open platform" IBM solution back to old fashioned POWER hardware because it was just better built and the "open" platform just resulted in finger pointing between BLADE, SAN, LAN, and OS people... And that was IBM support with their name on the box.. When it's truly mixed support it's going to be very terrible.
It's not like AIX or System i will ever get ported to non-IBM machines for better value. It's not like Apple is gonna dust off OSX and let me build racks of POWER servers either. Betting on "Linux" is a cop out. People use Linux on POWER when they have investments in POWER iron for their mainline business apps.. ROG, COBOL, etc.. And they don't want to pay AIX/enterprise licensing for web/email servers.
This is just the "last call" to show they tried to let the world love them more... And nobody will come. Then they close it down.
This is uncomfortably close to SPARC's story. Sun opened up SPARC designs in 2005. Four years later, the "Rock" project was canceled, marking an end for world-beating SPARC performance.
(Later Sun/Oracle chips have all been based on "Niagara," a low-end chip that didn't even hope to compete on performance. It was intended to be massively multicore and inexpensive, and it was at least one of those things.)
http://www.oracle.com/us/solutions/performance-scalability/s...
Also, if you look at the roadmap, you'll see there's even more performance improvements scheduled in the near future:
http://www.oracle.com/us/products/servers-storage/servers/sp...
On SPEC's general purpose benchmark, SPECcpu, Oracle published only three SPARC results. All three are from Fujitsu chips in re-branded Fujitsu systems. Oracle just doesn't publish results for their in-house chips. I imagine they ran the benchmarks and decided the poor results did not fit their marketing message.
https://blogs.oracle.com/BestPerf/entry/20130326_sparc_t5_sp...
http://www.spec.org/cpu2006/results/cpu2006.html
While I can't provide you my own detailed benchmarks, I can tell you anecdotally that the most recent SPARC hardware has significantly better single-threaded performance than the early T1, T2 series.
As an example, one of the projects I work on is written almost entirely in Python, and so is almost completely single-threaded (only the transport is multi-threaded using libcurl).
I had an opportunity to compare the performance on the T5 to the T1 and found it was practically night and day, and the performance of the program seems just as snappy as it does on any Xeon system I've used.
Anecdotally, I can tell you that given a choice between a fully-loaded Xeon box or one of the newest SPARC servers, most of the developers I work with will choose the SPARC server simply because builds take significantly less time.
Even the rate benchmarks are not very impressive. Oracle rigged the comparison in their press release. They compare a 16-core SPARC to an 8-core Intel. If we do apples-to-apples, the Oracle result is pretty humdrum:
- Oracle T5-1B, 16-core SPARC, 436 / 467
- Dell M520, 16-core Intel, 533 / 553
The M520 contains 2x E5-2450@2.1 GHz. This is far from the fastest Intel chip. You can get them up to 3.6 GHz. It's just a common and inexpensive configuration. Let's not even talk about the respective pricing.
Personally, I haven't used anything newer than a T1. Because I haven't seen anyone buy a new Sun box in that long!
- SPARC T5, 128 threads @ 3.6GHz, 463 / 467 -> 1.01 / 0.946 result/thread/GHz
- Intel Xeon E5-2690, 16 threads @ 2.9GHz, 357 / 343 -> 7.69 / 7.39 result/thread/GHz
Or to put it another way, the SPARC needs 800% more threads and a 24% higher clock to achieve only ~30% faster than the Intel. The POWER is somewhat, but not that much better, at 2.54 / 2.24.
I used this poor comparison because that was what was readily available in SPEC's published results for single-processor (SPECcpu) and parallel (SPECcpu_rate) benchmarks.
When I pointed out that results were provided for that benchmark, you then complained that it wasn't for the general benchmark but a portion of it.
You then proceeded to claim it wasn't an Apples-to-Apples comparison, but Oracle doesn't offer anything less than 16 cores for a T5. I don't think comparing products that don't exist is very useful so attempting to extrapolate what an 8-core version might be seems silly, especially since there's not a one-to-one correlation between cores and performance.
In addition to that, at last check, you can't pick the number of cores (precisely) that you want a processor to have when purchasing so it doesn't make any sense (in my personal opinion) to strictly compare core-to-core performance, since, as the other poster pointed out, core is essentially a definition at the whim of a vendor.
So, I'll just stick to refuting your original implication -- that SPARC isn't setting world-records anymore; it is in fact doing so. And in fact, SPARC continues to have far greater memory bandwidth, I/O bandwidth, and memory capacity compared to general Intel offerings.
So if you want to find out how fast a T5 will actually run your application, try one out and get real data instead of relying on benchmarks to make your decision. Personally, I think you'd be shocked at just how well most workloads perform if you actually tried a T5.
As to the price argument, the companies I've worked for or with in the past generally didn't care about that so much as the reliability of the system and it's capability. They have workloads that consume multiple terabytes of memory. They're using the servers to process transactions that are netting them millions of dollars with the servers they use, so saving a few thousand bucks doesn't matter to them.
In the end, you have to use the right tools for the right job. There's a reason that Oracle sells x86 servers too.
Personally, I don't care which architecture is being used as long as I get to use Solaris/ZFS.
You pasted an Oracle press release that focused on parallel performance.
I complained, and pointed out SPECcpu as a common measurement.
You responded with SPECcpu_rate, a different benchmark focused on parallel performance measurement. It is not surprising that Niagara-derivative chips do well. It is also not surprising that, core for core, they can't match modern commodity systems for density, performance, or cost.
p.s. The "reliability" argument went out the window with VMware. Every fortune 500 is using x86 with VMware HA to provide the redundancy that would have once come from enterprise RISC. (Given how few RISC systems were ever configured with redundant memory or CPUs, VMware-on-commodity is probably offering a substantially better service level.)
p.p.s. a basic 1U x86 server will typically have between 0.75 and 1.5 TB of RAM in it. Yes, TB, as in terabytes. Virtualization provides a market for compact, low-wattage systems with significant memory. That's just the 1Us. In 2014, "Large" x86 boxes are very large indeed.
Lastly, I, too, miss Solaris/SPARC. It was a great platform that I enjoyed working with. I don't miss the associated hardware support contracts that came floating by after the Oracle buyout. It has been several years since I found a contract where SPARC or Solaris were anything but legacy platforms. It's sad, but it's not a mystery.
The reliability argument doesn't "go out the window"; what do you think that software is virtualised on? And to top it off, you're losing a significant amount of performance by using a virtualisation solution like vmware.
In fact, there's entire segments of industry that won't accept the latency typical virtualisation technologies bring.
As for the 0.75 and 1.5 TB of ram argument, perhaps you missed what I said about terabytes. You have any Intel boxes with 32 terabytes lying around?
While you may have not seen any contracts floating around for SPARC hardware, I have quite recently..
Finally, as for "legacy" status, SPARC and Solaris have features not found anywhere else that are continuing to be developed and added even today. It's only a "legacy" platform if you completely ignore the technology there. And Solaris runs just fine on x86 thank you very much.
http://space.stackexchange.com/questions/729/why-did-the-esa...
The SPARC ISA is completely unencumbered as far as I know, and there are several copyleft implementations of the architecture that are synthesized into interesting products, often from overseas vendors (for example Navspark: https://www.indiegogo.com/projects/navspark-arduino-compatib... )
I ran (well, still do) a website dedicated to Sun/Solaris hardware and software users, was on the OpenSolaris external pre-release beta testing team/group, and heard lots of moaning about this, even from sales people within Oracle itself.
Pre-Oracle Sun was gracious and gave me a loaded SunFire T1000 to run the site and mailing lists, etc, on. Post-Sun Oracle gave me the finger.
Granted, that may seem like ancient history, but RDRAND is an indication that Intel still makes mistakes.
Less fragmentation is more efficient in many ways, but just like any monoculture, it is also fragile.
I first started using EMACS shortly after various IBM systems, and it's hard to express how obnoxious it was to have your keyboard lock while the computer was sending you stuff....
The question is serious because it's not clear to me the use cases of the zSeries including supporting this sort of thing (vs. editing your files on another computer like your PC before submitting them to the mainframe ... but I haven't touched anything in that domain since ... 1978 I think).
http://www-03.ibm.com/systems/z/os/zos/features/unix/library...
http://pic.dhe.ibm.com/infocenter/zos/v1r11/index.jsp?topic=...
Having said that, you can host Linux VMs under zVM. Binaries will use the zSeries ISA and the whole guest will run under the zVM environment. You can easily ssh to it. IIRC, Debian, Red Hat and SuSe support it.
This would probably be the least cost-effective way to edit text. On the other hand, few terminals were able to render monospaced text so beautifully.
However, I'm somewhat uncomfortable with Intel's increasing and complete near-domination of the "not-low-power" CPU market. I'd love to see other high-end silicon designers start to develop amd64-compatible processors that will encourage competition in the market. AMD's struggling with it, but there's plenty of scope for looking at interesting approaches to increasing performance or reducing power use.
Ahem, sorry to those suffering Mill Fatigue on HN, but any conversation about the ISA wars being over is a red rag to a bull with us! :)
Hope you like the talks on http://millcomputing.com/docs
NOTE: bulls react to movement not colour
I mean that the ISA wars are over in the sense of multiple relatively similar architectures competing against one another for negligible gain.
There's still plenty of scope for nontraditional architectures to come in and offer something different. We've already seen that in the more general availability of general-purpose GPU programming. As an architecture nerd, I'm already super-excited by Mill.
And their product plan was very conservative, with the "odd numbered" chips not being great advances on the even numbered ones, e.g. the 68030, which came out the same time as the first SPARC, used the 68020 microarchitecture with a 256 byte instruction cache and used the design shrink to put the MMU on the chip (but not the FPU). The 68040, three years later, wasn't a stunning success, was it? And the 88100, which came out a year after the 68020 certainly sucked up a lot of corporate oxygen, not to mention put into question the company's commitment to the 68K family.
And it's very dangerous for a company to depend on another company for one of the most critical components of its products, isn't it?
But, yeah, the economies of scale, the much lower unit sales to spread your Non-Recurring Engineering costs over, eventually doomed them to mostly bureaucratic and installed base niches when AMD successfully innovated for a short period and then Intel got its act together due to that threat.
But it's also in part 20/20 hindsight, e.g.:
Lots of people refused to believe that Moore's Law would last as long as it has; the corollary that you'd get higher speeds purely from design shrinks did end about a decade ago.
I'm not sure very many people "got" the Clayton Christensen The Innovator's Dilemma disruptive innovation thesis prior to his publishing the book in 1997. He really put it all together, how companies with initially cruddy products could in due course destroy you seemingly overnight.
In this case, how a manufacturer of rather awful CPUs (the 286 in particular, but the 8086/8 was no price except in cost; caveat, Intel support to people who design in their chips was stellar back then), could start getting its act together in a big way in 1985 with the 386, then seriously crack their CISC limitations with P6 microarchitecture (Pentium Pro), etc.
And note their RISC flirtation with the i860 in 1989, their Itanium debacle, etc. etc. More than a few companies would have committed suicide before swallowing their pride and adopting their downmarket, copycat's 64 bit macroarchitecture that competed with the official 64 bit one.
And that's not even getting into all the mistakes they made with memory systems, million part recalls on the eave of OEM shipments, etc. What allowed Intel to win? Superb manufacturing, and massive Wintel sales, I think.
As far as I can tell, Mill have realised that the real fight is for cache; the most expensive thing you can do on a modern CPU is a data stall. So the intent is to change the execution model enough to change data access patterns and make prefetch work properly, along with more efficient use of the cache.
The next most expensive thing is a branch mispredict, and Mill attacks that strongly as well.
And of course this is intimately dovetailed to cache, so Yes Yes Yes to everything you said too.
An IEEE I think article laid them out nicely into 4 categories:
Zero cost: what you put in microwaves, every cent counts.
Zero power: used to be obscure, the mobile market is of course making it very much less so, although those have elements of:
Zero time: speed is what counts, and plenty are willing to pay a hefty premium for it.
Zero units: say the military needs a CPU for a combat airplane which won't be made in more than 100s of units. Perhaps worth it for the prestige and getting the government to pay you to figure out neat things that might be usable in your bread and butter channels.
One modern example, but not with custom designs so much, is radiation hardened CPUs for space applications, a unit cost of say 100K is rather small in the bill of materials, the cost of failure in the high millions at minimum, could top a billion or billions, e.g. it would be very very bad if the computing subsystem(s) of the James Webb Space Telescope fail on its way to the Earth-Sun L2 point....
There's one strategy of increasing performance on a CISC that can't really be done with RISC - making existing complex instructions execute faster. This is particularly attractive since it means all existing software gets a boost without any software-level work. So in some sense, having complex instructions that would not be fast at the time the ISA was designed is like future-proofing. And x86/amd64 still has many places where this sort of optimisation can be done.
> I'd love to see other high-end silicon designers start to develop amd64-compatible processors that will encourage competition in the market.
Absolutely. It's unfortunate that many regard the ISA as being too complex, since it is within that complexity that is hidden a wealth of opportunities for optimisation.
You also had to worry about the ISA when you were doing something performance intensive (that's when you start taking advantage of multimedia instructions and so forth; ideally you have abstractions for this, but they'll only take you so far).
http://en.wikipedia.org/wiki/Intel#Market_share
In all honesty, beyond the cheap CPU market, most people have gone with Intel for a decade(s).
Intel has been dominating the Windows-compatible desktop computer space and recently made significant progress in the server space but if you lump together mobile and embedded processors and units sold, you'll see is nowhere near a leadership position in overall processor sales.
Latest production ARM in Nexus 5, 7 has 4 cores, integrate GPU, VPU (video processing unit.) Run from battery power. IMO, Intel is losing CPU war to ARM.
We are also seeing Qualcomm's 64 bits 8 cores going into the next gen phone soon.
ARM is a commodity market. Nvidia makes its money on GPUs and runs Tegra at a loss. ARM has a huge market cap but barely makes any profit. Samsung makes its money on the phone and other chips. TI found ARM so cutthroat that they've given up and quit. Qualcomm has an enormous market cap (larger than Intel's), but they make almost all of their money on the wireless side.
This is why ARM is so dangerous to Intel. While ARM moves upmarket and Intel moves downmarket, Intel is unable to compete on price because the competitors are operating so close to breakeven.
Has this changed?
In any case, TI makes almost all of its money on analog chips. "Embedded processing" was less than 8% of profits in Q1 2014. And that division includes microcontrollers.
And please don't downvote without fist spending ten seconds to look up the companies' financials.
I meant exactly what I wrote: "market cap." Qualcomm had a larger market cap that Intel (at least it did yesterday). It did not and does not have greater revenue than Intel, especially when you strip out the wireless side and only look at CPUs.
The point of bringing up market cap is to show how huge these companies are that Intel is competing with. Everyone knows Samsung is enormous, so I didn't bother pointing it out. A lot of people don't realize how huge Qualcomm has gotten. But it's not because of Snapdragon.
Unfortunately, the internal cost structure was such that the Ultra AXmp cost only a few percent less than a "proprietary" workstation/server from the parent company.
The moral of the story: sticking a high-margin, low-volume chip in a commodity board doesn't put it on a commodity cost basis.
In the end, it's the entire ecosystem... it's either affordable and accessible or it isn't. It's sad that IBM keeps coming out with pretty cool hardware (Cell, POWER, etc.) that nobody can afford to get convenient access to. :-(
And, IMO, when an architecture is that inaccessible to the hobbyists and experimenters and tinkerers, it really hinders it. I know it definitely diminishes my personal interest, and I want to be an IBM fan in many ways.
Ah, gotcha. Yeah, that's kind of a bummer if that is actually the case. I haven't really looked into buying any OCP servers, so I wasn't aware of that situation.
With or without them, it's probably too little too late for this architecture, sigh.
MIPS, UltraSparc, and Power, are sadly fading away.
Also wonder if its a year or two too late. AMD seems to have picked ARM as their plan B strategy realising that competing with Intel on x86 is hard to do with reasonable margins. Wonder if AMD would have picked POWER if it was available as an alternative a couple of years back.
ARM chips don't have to "beat" Intel. They just have to become good enough for desktop performance, while costing much less.
[1] - http://www.electronicsweekly.com/news/business/intel-lose-1b...
It was an outgrowth of his work developing the first optimizing compiler (Fortran) with Fran Allen which earned him a Turing Award. Good on ya, Johh, wherever and with whomever you might be sipping now. :-)
For example a Google plays around with racks of disposable X86 boards like candy. When their app becomes "fixed" rather than growing exponentially, they'll want to move to something like POWER because it's DESIGNED to work with dozens of CPUs sharing Petabytes of attached disk easily. Not the silly kludges like Blades, SANS or iSCSI or virtual machines people play with now to hide x86 OS vendor scaling limitations.
However: are there any inherent limitations in the x64 architecture which would make it impossible to achieve the same as with POWER, if you designed it from scratch for that kind of robustness?
I guess we are seeing another occurrence of Gresham's law applied to the technical field. "Bad architecture drives out good". In this case bad is cheaper, and this is the only thing we care about after all.
1. Open Firmware allows the system to load platform-independent drivers directly from the PCI card, improving compatibility. http://en.wikipedia.org/wiki/Open_Firmware
2. http://www.spscicomp.org/ScicomP16/presentations/Power7_Perf...
As far as I can tell the only reason EFI even exists is NIH syndrome. Intel could have just adopted OpenFirmware for x64, and they still should.
I'm not a hardware guy, but it always seemed like PowerPC chips underperformed others and ran as hot as hell while doing so.