First stage POWER9 Firefox JIT passes tests
talospace.com
talospace.com
Or IoT devices?
Or...
Or...
What detriment would that be?
https://www.nxp.com/products/processors-and-microcontrollers...
Is 3W is small enough for ya?
Semiconductor design teams don't exactly grow on trees, either. It was over a decade from Apple buying PA Semi whole (https://en.wikipedia.org/wiki/P.A._Semi ) to announcing the M1.
And of course if you're going to do that you already need the rest of the vertically integrated pipeline to build motherboards to put your chips on, peripheral IP to do all the other things other than processing, etc.
I haven’t checked, but I don’t expect Power to have good USB support, for example.
I also think Power targets a different performance range than ARM.
Reading https://en.wikipedia.org/wiki/OpenSPARC and https://www.oracle.com/servers/technologies/opensparc-overvi..., both of which mention no news after 2008, OpenSPARC looks on its deathbed to me.
I think people forget that the economies are different with hardware vs software; because you cannot eradicate the per-unit cost of hardware, paying a small part of that in license fees is not a big deal, especially since it comes with integration support that saves you a lot of non-recurring R&D expense. Whereas in software, being free makes it zero-friction and this has a huge impact on adoption.
Optimization is a bigger issue, though autovectorization in compilers is making it less of a problem than it used to be, as well as us nerd pioneers getting things upstreamed.
Is POWER9 compilation generally poor compared with other targets? It didn't seem so to me, apart from a pathological case on one benchmark set which IBM addressed swiftly. They've been supporting GCC for rather a long time.
I agree IBM's autovector stuff is quite good in gcc but there's no substitute for hand-rolled assembly sometimes.
It is valuable only if it has many users, e.g., application code, optimized for the ISA.
HW's job is to then run that code with good perf/cost.
OpenPOWER has little software. ARM has a lot of software.
So from that POV, ARM is already many orders of magnitude more valuable than OpenPOWER.
But it doesn't end there. Do you need some software to be extremely optimized for ARM? ARM can do this for you at resonable price, no need to hire.
Also, for OpenPOWER, you need to hire 50-100 Facebook engineers, at 400k$/year, and it'll take them >3 years to produce a chip design, which then needs to be verified, etc. and then needs to be built, so you'll need a fab, specialized on OpenPOWER, or not. A fab churns 40k chips/month, so how many chips / month do these companies need ?
With ARM, you pick one of the many ARM farms, and there is little for you to do. And you get 5 engineers, and they just customize an ARM design to your needs. And they ship in 1 year instead of 3. And next year ARM gives you a way to update your chip to the next generation. And if next year you need some other feature, ARM gives it to you. And if you need software, like C library, profilers, math, all that is supplied by Arm.
And they take royalties on chips you built, and.... and....
So ARM is many orders of magnitude cheaper in perf / $ than OpenPower. Not only is the hardware better, but it is better, cheaper, has more software, and tools, and teams of experts ready to help your team, etc.
Porting most software to ARM64, Power, or RISC-V involves typing some variation of "make." Only a small percentage of software written in C/C++ or ASM is problematic. Anything in a higher level language like Go or a newer language like Rust is generally 100% portable.
Switching from X86_64 to ARM64 (M1) for my desktop dev system was trivial.
Endian-ness used to bite, but today virtually everything is little-endian above embedded. Power and some ARM support both modes but almost always run in little-endian mode (e.g. ppc64le).
- Have you ever multiplied a matrix with a vector, or a matrix with a matrix (GEMM) using BLAS?
- Have you ever done an FFT ?
- Have you used C++ barriers? Or pthreads? Or mutexes?
An optimized implementation achieves ~100% of theoretical peak performance of a CPU on all of those, and these are all tailored to each CPU model.
There is software on any running system doing those things all the time. Running at 0% of the peak just means increased power consumption, latency, time to finish, etc.
Generic versions perform at < 100%, often at ~0% (0.1%, 0.001%, etc.) of theoretical peak.
Somebody has to write software for doing this things for the actual hardware, so that you can then call them from python.
IBM has dozens of "open source" bounties open for PowerPC, and they pay real $$$, but nobody implements them.
---
Porting software to PowerPC is only as simple as doing make if the libraries your software uses (the C standard library, the libm library, BLAS, etc. ) all have optimized implementations, which isn't the case.
So when considering PowerPC, you have to divide the paper numbers by 100 if you want to get the actual numbers normal code recompiled with make gets in practice. And then you have to invest extra $$$ into improving that software, cause nobody will do it for you.
While it isn't necessarily clear what peak performance means, MKL or OpenBLAS, for instance, is only ~100% of serial peak on large DGEMM for a value of 100 = 90; ESSL is similar. I haven't measured GEMV (ultimately memory-bound), but I got ~75% of hand-optimized DGEMM performance on Haswell with pure C, and I'd expect similar on POWER if I measured. Those orders of magnitude are orders off, even for, say, reference BLAS. I don't know why I need Python, but the software clearly exists -- all those things and more (like vectorized libm). You can even compile assorted x86 intrinsics on POWER, though I don't know how well they perform relative to on equivalent x86, but I think you're typically better off with an optimizing compiler anyway.
I've packaged a lot of HPC/research software, which is almost all available for ppc64le; the only things missing are dmtcp, proot, and libxsmm (if libsmm isn't good enough).
If you have used Sierra since the beginning, we have seen significant performance increases over the years, because the people using it have actually been discovering and then either getting IBM to fix, or fixing themselves, most of the software.
Compared with Power 10, I'd say that Power 9 is "mainstream" (many clusters available), and from the Power 9 CPUs in existence, IBM's are the most mainstream of them all.
Take the Power 10 ISA, build your own CPU that significantly differs from IBM's, and good luck with optimizing all the software above. It can be done, and dumping it on a couple of HPC sites where then scientists and staff won't have any change but to use it for 4-6 years is a good way to get that done.
But for a private company that just wants to deliver value, ARM is just a much better deal, cause it saves them from having to do any of this.
Most Linux software has being ported over.
Works efficiently is step 10000.
x86 is at step 10000, ARM at step 5000, power is at step 0.
Firefox "worked" before this post on power. Now somebody put enough effort to actually make it usable.
The fact that you don't see people complaining about Firefox PowerPC performance on Linux is not because performance was good - it was unusably slow - but because nobody uses Firefox on Power.
Think about what that means. Think about how many bugs in Firefox are reported _every day_ for x86 and ARM, and how many are reported for PowerPC. Is that also because the PowerPC version has no bugs? (no, it is because nobody uses it, nobody reports them, and nobody fixes them).
I agree with your general point, but I do believe that Power is the most "practical" ISA after x86 and ARM - albeit it's a distant third, it's definitely not at 0. It has the full support of a bunch of mainstream distros, public container registries have a decent amount of support for their images, and people actually run pretty serious workloads on Linux on Power.
Power does have a lot of niche backing, albeit it's continuously being hurt by IBM's total lack of interest in doing anything but push it beyond the billion dollar contracts they're milking with it. That's totally destroying any mindshare Power has. There's really no way to get a cloud shell on a modern Power machine, or physical access to a modern one without forking over thousands of dollars for the privilege (the latter only really is possible due to Talos' amazing efforts, bless em).
Afaik this article is a big deal because it's JIT, so a big chunk of architecture specific code to get good performance. But most code is not going to be that.
That's not to say that architecture specific bugs will not exist. But i think your outlook on this is a little pessimistic.
Also, ARM has always (?) been about getting the most out of limited resources, whereas POWER is about performance at any cost. With modern ARM designs, the performance is getting close to, and even exceeding, that of traditional desktop and server CPUs, while still being frugal with resources. POWER is still, well, power-hungry.
The biggest single cost for AWS etc. is per-rack running costs, encompassing power, cooling etc. It's hard to overemphasise just how much this dwarfs all other costs. To optimise those costs you've got to cut down the power consumption and associated heat production.
And finally, what benefits does POWER and SPARC brings to the table? The licensing cost from ARM is tiny in the grand scheme of things. I like open ISAs like POWER and OpenSPARC, but from a business POV it just doesn't make any sense.
Radiation hardened implementations, e.g. spacecraft, is one niche, though I don't know whether that has anything particular to do with the ISA.
Can change the endianess at run time.
Their new cache architecture is very unusual. A CPU can use the cache from another chip. Data is stored encrypted.
;)
From an efficiency perspective does that mean the chip has to do both so worst of both worlds?
[I dont know much about the world of cpu design, these might be stupid questions]
The POWER ISA was used in PowerPC which was used for the successors of a few 68k machines (most famously the Macintosh) and in that case the OS was built for big-endian. So having big-endian support was key there.
As for endian shifts, technically every OpenPOWER chip goes big for every OPAL call into the low-level HAL, even if the OS is little. The overhead is minimal. I can't think of much application use for that, though (per-page endianness which some PowerPCs supported is much more useful).
The CPU probably works on just one endianess and convert the data format when reading from memory. The overhead is on kepping track when to do it. But Im speculating, havent looked into this.
The IBM Power is one implementation of the Power ISA.
The chips used on the mars rovers and nintendo Wii as ppc as well.
Yeah, Cortex and NVIDIA cores do (and probably quite some others).
However, Apple-designed recent cores don’t implement big endian support at all.
Is OpenPOWER trying to compete with RISC-V? Whats the benefit compared to modern ARM or x86_64 silicon & ISAs?
All this with fully open firmware, and an open ISA (as of the last couple years). The CPU implementation itself is not open, but all firmware and procedures for initializing the CPU are open. For people interested in that sort of thing, it's appealing as a practical computer with a full PCIe implementation with actually decent performance, compared to essentially every other open source platform.
What they're really missing is a midrange product for a midrange price. I can't blame them for avoiding the low end, but can't I get anything for less than $2000?
This seems to be universally true for all kinds of UNIX workstations and servers.
Repeat until the hardware is rare and worth something?
So while it's unfortunate, it's not a case of ignoring the low end deliberately but mostly flows from the economic realities of not having anywhere near the addressable market of x86 or ARM. The small community of ppc64(le) enthusiasts is very much hoping for a future where this changes, however small that chance might be...
IIRC some of the latest AMD boards to end up with coreboot support is using opterons from 2011-2013.
For what it's worth:
$ lscpu | head -5
Architecture: ppc64le
Byte Order: Little Endian
CPU(s): 160
On-line CPU(s) list: 0-159
Thread(s) per core: 4From my point of view it is interesting because it provides an alternative to x86 and ARM. The more competition, the better. Competition will push the hardware forward. Imagine we only had only one CPU architecture and one CPU maker.
We have far less CPU architectures than we used to have.
These devices support basically all the DRM GPU drivers in Linux, and when coupled with Mesa, have very fast and responsive GUIs. Both Firefox and Chromium run well, but up until now, Firefox has been using a pure interpreter to run javascript. This is fine for sites not using much javascript, but a big chunk of the modern web quite literally loads megabytes worth of JS on a page, and it can really chug under the interpreter.
So it's pretty exciting that we'll have a second browser with a proper JS engine, Chromium being the first (IBM ported V8 to ppc64le, for node.js, but the port works for running chromium as well)
Also because it may not be clear, the other big arches all already have JIT compilers. x86, amd64, 32 and 64 bit ARM, (and I think MIPS does, on either chromium or firefox, don't recall), so this is less about boosting performance, and more about reaching "baseline expected performance".
Anyway, I'm trying to go as fast as I can to get an actual browser mounted. But passing the test suites in totality, run two different ways, suggests a high probability of success at this point.
(A) at the very least my every attempt to buy a POWER9 system has been thwarted, mainly by the manufacturer themselves constraining availability or otherwise being unable to supply, and
(B) POWER10 has IP issues that have made it unattractive to Raptor CS, and
(C) they have in the past made noises about it not being worth the effort to continue to sell POWER9 systems to the public because of the support overhead
...I have to ask if this is effort well-spent, or if the sweat would be better poured into something with a more certain future.
Not that RISC-V boards currently are anywhere near fast enough for daily driver use, but I'm leaning in that direction.