IBM’s 24-core Power9 chip
nextplatform.com
nextplatform.com
seriously? love to see how/why that was put in the die
At it's core a `SELECT FROM` statement is a filter lambda (without any joins, or before/after a join has been done depending on query optimization).
Fundamentally this is because an SQL table is a flat array. With row's being flat array's within the table's memory space.
Calculating the pointer to fields within rows within a table is ridiculously parallel. You're limited only by SIMD size. (this is 1 FMA + 1 add against a constant ~2 ticks).
Also custom wide SIMD registers can allow for `VARCHAR(255)` equality statements to be done in 1 processor cycle.
I would love to see some benchmarks of how much this speeds up some sample workload queries, if you happen to know of some (preferably not from Oracle's marketing dept.)?
The only ones you'll find by 3rd party are non-direct comparisons like transactions per second in Oracle-SQL. Or Core Count.
But they'll still cheat us on small things. We request information on a training seminar (we had a free attendance voucher), and they lose our account information until 1 week after the seminar ended.
It's to the point where I just laugh. Everyone knows their cheating us, but upper management won't change.
I mean I'd rather use Postgres. But large corps like having a vendor to call.
http://www.enterprisedb.com/products-services-training/produ...
Alas.
Good to see new developments in this line. The more (open) hardware platforms the better, IMO.
They killed the desktops. That, essentially, made SPARC mostly invisible to most people. They also do a crappy job of reminding the world that a software company (as they position themselves) can make decent hardware.
Edit: Found this in PowerISA_V3.0.pdf
Load Doubleword Monitored Instruction: Specifies a new instruction and facility that improves the performance of the garbage collection process and eliminates the need to pause all applications during collection.
The ldmx instruction loads the same doubleword in storage as the ldx instruction.
ldmx is intended for use by applications when loading data objects during times when objects are being moved as part of a garbage collection process. In this type of usage, all loads of object pointers by applications are performed using ldmx. Whenever objects within a given region of memory are to be moved, the garbage collection program sets the Load Monitored Region registers to encompass the region being moved, and sets BESCRGE LME to 0b11 to enable Load Monitored event-based branches.
...
LoPAPR http://openpowerfoundation.org/?resource_lib=linux-on-power-... IODA2 http://openpowerfoundation.org/?resource_lib=openpower-io-de...
(although i'm 99% sure the IODA doc is missing a heap of stuff).
>The Load Monitored Region registers are used to specify a Load Monitored region, and to specify the sections of the region that are enabled. The Load Monitored region is a contiguous region of storage specified by the Load Monitored Region register. Load Monitored regions range in size from 32 MB to 1 TB. All regions are powers of 2 in length and are aligned (see Section 1.11.1). The Load Monitored region is divided into 64 sections of equal length, each of which can be separately enabled. The Load Monitored Section Enable Register specifies which sections are enabled.
>Load Doubleword Monitored Instruction: Specifies a new instruction and facility that improves the performance of the garbage collection process and eliminates the need to pause all applications during collection.
Programming Note:
>The ldmx instruction loads the same doubleword in storage as the ldx instruction.
>Idmx is intended for use by applications when loading data objects during times when objects are being moved as part of a garbage collection process. In this type of usage, all loads of object pointers by applications are performed using ldmx. Whenever objects within a given region of memory are to be moved, the garbage collection program sets the Load Monitored Region registers to encompass the region being moved, and sets BESCR GE LME to 0b11 to enable Load Monitored event-based branches.
>Subsequently, if an application program loads (ldmx) a pointer into an enabled section of the Load Monitored region, an event-based branch (EBB) will occur. The EBB handler will load the instruction from storage, and decode the instruction to determine the effective address from which the pointer was loaded. The handler will then load the pointer (obtaining the same value the application obtained), and determine where the corresponding object has been moved to, and update the pointer in storage so that it points to the object's new location. The EBB handler may also take other actions such as updating additional pointers, depending on the situation. After this processing is complete, the EBB handler sets BESCR LME LMEO to (1 0) since the taking of the EBB set these bits to (0,1), and then executes rfebb 1 to re-enable EBBs and return to the application program at the ldmx. The application program re-executes the ldmx, which now returns the updated pointer (which no longer points into an enabled section of the Load Monitored region), and continues.
>Other variations and extensions of the above pro- cedure are also possible.
https://en.wikipedia.org/wiki/Intel_iAPX_432#Garbage_collect...
1: http://www.nextplatform.com/2016/04/06/inside-future-google-...
2: http://www.pcworld.com/article/3053092/ibms-power-chips-hit-...
- Linus Torvalds, 2009 (http://yarchive.net/comp/linux/coding_style.html)
* Peek a value that's MMIO
* EIEIO
* Poke right back into MMIO
* EIEIO
[1] http://www.ibm.com/support/knowledgecenter/ssw_aix_61/com.ib...
I imagine it shouldn't be too difficult (relatively speaking) to do the port as a result.
The GOARCH for power is ppc64/ppc64le AFAIK.
(s390x is the mainframe processor, aka z/Architecture, the latest being the z13 which is most definitely not a power8/9.)
It seems the vast majority of people who consider themselves knowledgeable about computers are just massively ignorant/confused by IBM's product lines and capabilities (and don't even get me started about the nonstop and openvms). Which is a bit of a shame, particularly in the case of IBM i, which is probably one of the most interesting OS's still in support/production.
edit: failed to refresh
As others have said s390x is nothing to do with Power.
I'm not sure if rust targets that platform/arch combo, but I could see it being deferred while they focus on more common ones.
Edit: It just occurred to me that Node's (V8's) optimized JIT probably uses ASM or at least routines optimized for x86, so that would complicate deployment.
1: https://nodejs.org/en/download/package-manager/#freebsd-and-...
> I'm not sure if rust targets that platform/arch combo
It's the PPC bit that's tough at the moment, we do provide tier 2 support for {free,net}BSD on x86_64.I definitely built and ran Erlang on a G4 powerbook several years back.
Node also supports powerpc. Support was added courtesy the team at IBM.
I have both working on my PowerMac G5, running FreeBSD 10.2.
That said, they're expensive like any high end datacenter-grade hardware.
More like a minicomputer or mainframe level power?
FreeBSD runs on modern Power I believe, but the other BSDs only run on older hardware AFAIK, although IBM is pretty good at providing hardware so that may have changed.
I am not against IBM POWER CPUs, I was just disappointed by POWER8...
Also, I have a major sad for our industry when people struggle so much with alternative architectures. Tons of software absolutely should not care including JVM languages, scripting languages, Golang. Most userland compiled languages should also minimally care. For lower level C stuff, regular compilation on non-x86 archs is often directly reflected on overall code quality (new compiler warnings, things like alignment correctness, cache/memory coherency model assumptions, data type assumptions, etc). It is somewhat hard to port operating systems and "efficiency libraries" that invoke platform features, but it also takes a comparatively small number of people to do.
ME, and the absolute refusal of Intel to allow its own customers to disable it, is very scary to the tinfoil hat crowd.
Unfortunately, Raptor Engineering is still early in development of their POWER8 board, and even that may ultimately turn out to be vaporware.
If they wait another 6 months it simply won't be economical to release it ever.
Power8/9 are very decent processors with a lot of raw power, but it will take a more concentrated and long term effort to have a level playing field in the software landscape.
I've developed for and used Linux and BSD on several platforms. Moving to a different one is trivial unless you're dependent upon proprietary binaries which can't be rebuilt. If you're using open code, or code which you have the ability to rebuild for the new platform, you're not locked in.
(In the '90s, the PowerPC had very competitive price/performance, but there wasn't a viable desktop OS available after Apple terminated the original MacOS clone program. 5-10 years later Linux would have been much more ready, but the chips were not competitive with Intel anymore.)
Their first reminder was perhaps netbooks, using semi-retired Celerons to make small and cheap clamshell web browsers and media players.
Their second, and ongoing, reminder is smartphones/tablets.
Without an operating system, IBM and Motorola didn't have an incentive to build PC chips anymore. The G4 happened because it was already in development and the vector extensions were useful for Motorola's embedded ambitions. The G5 happened because Apple basically paid IBM to make a desktop chip out of their POWER designs, AFAIK.
Around 2004, there was a startup PowerPC maker called P.A. Semi [1] that apparently competed for Apple's Mac CPU business. After Apple went Intel, they acquired P.A. Semi to design iPhone chips instead.
Here is a random google search for spec performance over time of assorted CPU arches.
http://preshing.com/20120208/a-look-back-at-single-threaded-...
The only obvious case in that graph where PPC>x86 is in mid 1995.
I can't find the one I saw a couple years ago detailing the core/clockrate comparisons of PPC's shipped in Macs with common PCs, that was much easier to read and far more obvious.
http://www.nxp.com/products/microcontrollers-and-processors/...
They're also deployed on FPGA's as embedded CPU's.
(full disclosure - I worked on tiny pieces of the code and stuff adjacent to it in the design)
since you seem to know about the power consumption, can you please share some numbers about the power consumption of the jcore cpu (on a comparable process to other cpu's build for IoT devices, usb controllers etc)?
jcore is not intended to be a replacement for the main cpu, so this would be really interesting to me, and would save me from asking on the mailing list. ;P
[1] http://se-instruments.com/press_releases/sei-adopts-open-sou...
https://github.com/ARM-software/arm-trusted-firmware http://www.tianocore.org/edk2/ (for more platform support add git.linaro.org/uefi/OpenPlatformPkg.git)
http://www.nextplatform.com/wp-content/uploads/2016/08/ibm-h...
Edit: This one, I think: https://upload.wikimedia.org/wikipedia/commons/4/44/Intel_80... It's nice that you can almost see individual wires.
My dad got it while working at the university and at some point in time in high school gave it to me. The Lucite is pretty scratched up by now, but it has a weird sentimental value to me so I'm planning on getting it buffed out.
[0] http://www.extremetech.com/wp-content/uploads/2014/04/ibm-po...
[1] http://images.anandtech.com/galleries/693/DieShot_GF100_Arch...
[2] (forum thread with many die shots) http://www.cpu-world.com/forum/viewtopic.php?t=19888
There are some photos from it at http://www.wired.com/2007/07/core-memory/
Is this a new ISA or just the latest revision in their existing ISA?
Also, what's the relationship with Freescale (NXP) here, who sells pretty decent "PowerPC" chips for server-ish platforms, such as the T2080.
Do they license the ISA from IBM sort of like how ARM works?
This is going to be the latest refinement of the POWER ISA, replacing the current POWER8 used by IBM, and licensed by some Chinese company for not yet released products.
FreeScale/NXP are licensing PowerPC.
The Freescale part of NXP was formed out of Motorola's semiconductor division - Motorola, in turn, was part of the original AIM Alliance with Apple and IBM back in the early 90s that developed PowerPC. So they've been making PowerPC stuff since the very beginning. I would assume they have a license of some sort from IBM.
[Disclaimer: IBMer, opinions my own]
But until Moore's Law is really buried and forgotten, new architectures won't be able to compete on already filled niches.
The actual difference is Intel has proprietary fab technology one or two generations ahead of TSMC etc.
[1] The actual numbers for 2014 are: Intel shipped 400 million chips, ARM partners shipped 12 billion chips.
That's what I said.
Sometimes weird stuff like supporting esoteric platforms happens because the companies in question are so large, that it makes sense to spend a few $100k on that project to make Intel squirm and get better pricing from Intel on your bulk deal. Keep in mind, Microsoft have their own cloud service...
There are a ton of Docker images that are not going to work on anything other than x64.
Lack of binary compatibility is a huge hurdle to overcome for any new CPU architecture.
x64 has the same issue in the mobile space against ARM.
Also, IIRC they've had POWER servers available for quite some time (possibly already POWER7), so maybe there is demand. Assuming so, it is either actually more efficient to a great degree than x86-64, or sufficiently more efficient so that it can solve some problems the x86-64 ones can't at all.
[1] https://www.online.net/en/dedicated-server/dedibox-power8
I'd love to see some market forces nudging Intel to offer CPU's and Motherboards without ME.
and: https://www.raptorengineering.com/TALOS/prerelease.php
Supposedly this could be a way out of Intel ME. Gotta see it first though, but I have one pre-ordered (whatever that means).
edit: should have read till the end.
Like an Apple.com/iphone page for Power systems :-)
Search for Power9 on IBM.com yields practically nothing.
http://www.ibm.com/search/esas/search?q=power9&v=17&sn=23&us...
Perhaps the website is all maintained by the folks busy in "the cloud".
Today's cutting edge fabrication of multi-core chips makes it quite unlikely to get many dies to manufacture where all the CPU's are working at high speeds. Chips must run at a single frequency, so the slowest frequency is the bottleneck. If you fabricated 32 cores, and 1 fails testing, the chip is garbage.
So the strategy is that the chips are fabricated with 32 cores, and then they speed-test the cores and take fastest 24 cores. The remaining 8 are fuse-disabled, and the end user sees a 24 core chip.
Note this strategy also allows them to produce a 16-core chip with the same die, if for example only 20 cores work at high speeds.
So you get both high yield and high speed chips.
[1] https://news.ycombinator.com/item?id=12351729
http://www.nextplatform.com/wp-content/uploads/2016/08/ibm-h...
We had to test against the OSs we'd since dropped support for, which included HP-UX, AIX and s390.
The first time I heard someone mention a "31bit" build, I thought they were taking the piss :P It just seemed so odd at the time...
Open source software has been a powerful economic force: the fact that so much software exists that's designed to target many similar platforms and is frequently recompiled for those targets means that the impact to the end-consumer of switching among x86/ppc/arm is much lower than it was in the past.
For instance, in the VPS hosting market, you see most providers offer "vCPUs" (virtual CPUs), because then they can sell their customers "four cores" (or 20 cores) on the cheap, and so on. It's a good marketing strategy.
What if they can offer real cores for the same low price? That would be pretty appealing to their customers. They can also offer "24-core dedicated servers" on the cheap, and so on.
Also, in the long run, it would be best for hosting and cloud providers if they would at least adopt a strategy of 50% Intel chips, and 50% AMD + Power + ARM. That should get them much better prices in the future (from all the chip providers), if they did that. Monopolies don't serve anyone but the monopolist.
Highly recommended. http://millcomputing.com/docs/
However, we'll see what happens if and when they produce real silicon whether the concept will be another Itanic, even AMD moved away from VLIW designs for their GPU's for performance reasons ~4 years(though, to be fair, they don't exactly have the $$$ to pay for tremendous amounts of driver developers to optimize their code generation anymore).