Next-Generation IBM POWER10 Processor
newsroom.ibm.com
newsroom.ibm.com
- They leapfrogged everyone else with PCIe v5 and DDR5
- 1 TB/s memory bandwidth, which is comparable to high-end NVIDIA GPUs, but for CPUs
- Socket-to-socket interconnect is 1 TB/s also.
- 120 GB/s/core L3 cache read rate sustained.
- Floating point rate comparable to GPUs
- 8-way SMT makes this into a hybrid between a CPU and a GPU in terms of the latency hiding and memory management, but programmable exactly like a full CPU, without the limitations of a GPU.
- Memory disaggregation similar to how most modern enterprise architectures separate disk from compute. You can have memory-less compute nodes talking to a central memory node!
- 16-socket glueless servers
- Has instructions for accelerating gzip.
Where do you get the FP performance, exactly, and for what value of "FP"? It's unclear to me in the slides from El Reg, which appear to be about the MMA specifically, and it's not clear what the SIMD units actually are. (I don't know if that's specified by the ISA.)
One thing is that it presumably has a better chance of keeping the SIMD fed than some.
Has there been any research for finding out how much these custom instructions have accelerated modern computational power in practice?
There are a lot of these instructions in modern CPUs for AES, gzip, SSL, etc but you have to have a library coded to support them.
I'm curious what the performance improvements look like in practice.
I recently did a bunch of tests to see what the "ultimate bottlenecks" are for basic web applications. Think latency to the database and AES256 throughput.
Some rough numbers:
- Local latency to SQL Server from ASP.NET is about 150 μs, or about 6000 synchronous queries per second, max.
- Even with SR-IOV and Mellanox adapters, that rises to 250 μs if a physical network hop is involved.
- Typical networks have a latency floor around 500-600 μs, and it's not uncommon to see 1.3 ms VM-to-VM. Now we're down to 800 queries per second!
- Similarly, older CPUs struggle to exceed 250 MB/s/core (2 Gbps) for AES256, which is the fundamental limit to HTTPS throughput for a single client.
- Newer CPUs, e.g.: AMD EPYC or any recent Intel Xeon can do about 1 GB/s/core, but I haven't seen any CPUs that significantly exceed that. That's not even 10 Gbps. If you have a high-spec cloud VM with 40 or 50 Gbps NICs, there is no way a single HTTPS stream can saturate that link. You have to parallelise somehow to get the full throughout (or drop encryption.)
- HTTPS accelerators such as F5 BIG IP or Citrix ADC (NetScaler) are actually HTTPS decelerators for individual users, because even hardware models with SSL offload cards can't keep up with the 1 GB/s from a modern CPU. Their SSL cards are designed for improving the aggregate bandwidth of hundreds of simultaneous streams, and don't do well at all for a single stream, or even a couple of concurrent streams. This matters when "end-to-end encryption" is mandated, because back end connections are often pooled. So you end up with N users being muxed onto just one back-end connection which then becomes the bottleneck.
Modern processors can do ~2.1GB/s per core for AES-256-GCM (about 2x what they can do for AES-CBC).
Both EPYC and Xeons have 2-AES units per core now. But CBC can only effectively use one of them at a time. (Block(n+1) cannot be computed until block(n) is done computing. Because Block(n) is used as input into Block(n+1) in CBC mode).
AES-256-GCM can compute block(n) and block(n+1) simultaneously. So you need to use such a parallel algorithm if you actually want to use the 2x AES pipelines on EPYC or Xeon.
Of course, the underlying TCP latency is significantly lower. Using Microsoft's "Latte.exe" testing tool, I saw ~50 μs in Azure with "Accelerated Networking" enabled. As far as I know, they use Mellanox adapters.
Something I found curious is that no matter what I did, the local latency wouldn't go below about 125 μs. Neither shared memory nor named pipes had any benefit. This is on a 4 GHz computer, so in practice this is the "ultimate latency limit" for SQL Server, unless Intel and AMD start up the megahertz war again...
It would be an interesting exercise comparing the various database engines to see what their latency overheads are, and what their response time is to trivial queries such as selecting a single row given a key.
Unfortunately, due to the DeWitt clauses in EULAs, this would be risky to publish...
Very often, you'll see n-tier applications where for some reason (typically load-balancers), the requests are muxed into a single TCP stream. In the past, this improved efficiency by eliminating the per-connection overhead.
Now, with high core counts and high bandwidths, some parallelism is absolutely required to even begin to approach the performance ceiling. If the application is naturally single-threaded, such as some ETL jobs, these single-stream bottlenecks are very difficult to overcome.
In the field, I very often see 8-core VMs with a "suspicious" utilisation graph flatlining at 12.5% because of issues like this. It boils my blood when people say that this is perfectly fine, because clearly the server has "adequate capacity". In reality, there's a problem, and the server is at 100% of its capacity an the other 7 cores are just heating the data centre air.
It just doesn't make sense for the HW engineers and chip designers to optimize backwards for older software algorithm instead of optimizing forward for newer algorithms.
I recently benchmarked it versus the ancient "DEFLATE" algorithm, as seen in the .NET Framework.
They're virtually identical in speed using typical settings.
Zstandard only has a performance advantage if using the prepared dictionaries, in which case it is genuinely faster.
If you don't go to the effort of "training" a dictionary and using it, there's zero benefit to Zstandard.
GPUs will always have more GFlops and memory bandwidth at a lower cost. They're specifically built GFlop and memory-bandwidth machines. Case in point: the NVidia 2070 Super is 8 TFlops of compute at $400, a tiny fraction of what this POWER10 will cost.
If POWER10 costs anything like POWER9, we're looking at well over $2000 for the bigger chips and $1000 for reasonable multisocket motherboards. And holy moly: 602mm^2 at 7nm is going to be EXPENSIVE. EDIT: I'm only calculating ~2 TFlops from the hypothetical 60x SMT4 Power10 at 4GHz. That's no where close to GPU-level Flops.
However, the CPU-GPU link is slow in comparison to CPU-L1 cache (or even CPU-DDR4 / DDR5). A CPU can "win the race" by using SIMD onboard, completing your task before it even spent the ~5-microseconds needed to communicate to the GPU.
----------
With that being said: POWER10 also implements PCIe 5.0, which means it will be one of the fastest processors for communicating with future GPUs.
A64FX (in Fugaku, the current #1 machine on all popular supercomputing benchmarks) has shown that CPUs can compete with top-shelf GPUs on bandwidth and floating point energy efficiency.
Summit has 4,608 nodes x 6 GPUs each, or 27,648 V100 GPUs. It also was built back in 2018.
---------
While Fugaku is certainly an interesting design, it seems inevitable that a modern GPU (say A100 Amperes) would crush it in FLOPs. Really, Fugaku's most interesting point is its high rate of HPCG, showing that its interconnect is hugely efficient.
Per-node, Fugaku is weaker. They built an amazing interconnect to compensate for that weakness. Fugaku also is an HBM-based computer, meaning you cannot easily add or remove RAM (like a CPU / GPU team can configure to more, or less RAM by adding sticks).
These are the little differences that make a difference in practicality. But yes, A64FX is certainly an accomplishment, but I wouldn't go so far as to say its proven that CPUs can keep up with GPUs in terms of raw FLOPs.
HPCG mostly tests memory bandwidth rather than interconnect, but Fugaku does have a great network.
Adding DRAM to a GPU-heavy machine has limited benefit due to the relatively low bandwidth to the device. They're effectively both HBM machines if you need the ~TB bandwidth per device (or per socket).
Normalizing per node (versus per energy or cost) isn't particularly useful unless your software doesn't work well with distributed memory.
This POWER10 chip under discussion has 1TB bandwidth to devices with expandable RAM.
Yeah, I didn't think it was possible. But... congrats to IBM for getting this done. Within the context of this hypothetical POWER10, 1TB bandwidth interconnects to expandable RAM is on the table.
edit: also, vocals are most latency-sensitive if you are the one singing and hearing it played back at the same time.
GPU latency is ~5 microseconds per kernel, plus the time it takes to transfer data into and out of the GPU. (~15GBps on PCIe 3.0 x16 lanes). Given that audio is probably less than 1MBps, PCIe bandwidth won't be an issue at all.
Any audio system based on USB controls would be on the order of 1000 microseconds of latency (1,000,000 microseconds/1000 USB updates per second == 1000 microseconds). Lets assume that we have a hard realtime cutoff of 1000 microseconds.
While GPU-latency is an issue for tiny compute tasks, I don't think its actually big enough to make a huge difference for audio applications (which usually use standard USB controllers at 1ms specified latency).
I mean, you only have room for 200 CPU-GPU transfers (5-microseconds per CPU-GPU message), but if the entire audio-calculation was completed inside the GPU, you wouldn't need any more than 1-message there and 1-message back.
That would indicate their theoretical peak performance is higher. Unless you can line up your data so that the GPU will be processing it all the time in all its compute units without any memory latency, you won't get the theoretical peak. In those cases, it's perfectly possible a beefy CPU will be able to out-supercompute a GPU-based machine. It's just that some problems are more amenable to some architectures.
For games and graphics? I don't think so, a GPU has a ton of dedicated hardware that is very costly to simulate in software: triangle setup, rasterizers, tessellation units, texture mapping units, ROPs... and now even raytracing units.
Many ARM devices boot from GPU.
> - Floating point rate comparable to GPUs
No where close. POWER10 caps out at 60x SMT4, and only has 2x 128-bit SIMD units per SMT4. That's 480 FLOPs per clock cycle. At 4GHz, that's only 1.9 TFlops single-precision compute.
An NVidia 2070 Super ($400 consumer GPU) hits 8.2 TFlops with 448 GB/s bandwidth.
> - 8-way SMT makes this into a hybrid between a CPU and a GPU in terms of the latency hiding and memory management, but programmable exactly like a full CPU, without the limitations of a GPU.
Note that 8-way SMT is the "big core", and I'm almost convinced that 8-SMT is more about licensing than actually scaling. By supporting 8 threads per core (and doubling the size of the core), you get a 30-core system that supports 240-threads. That's a licensing hack: since most enterprise software is paid per-core.
I'd expect that 4-SMT will be more popular with consumers (ie: Talos II sized systems). Similar to how Power9 was
I don't think that's entirely fair: Massively multi-user transactional applications (read: databases) are right in the wheelhouse of POWER, and they're exactly the kind of applications that benefit most from SMT. Lots of opportunities for latency hiding as you're chasing pointers from indexes to data blocks.
Shit. That's actually awesome and could easily pay for itself.
NVidia 2070 Super also has FP16 matrix-multiplication units, achieving 57 FP16-matrix TFLOPs. These are the NVidia "Tensor Cores".
Ampere (rumored to be released within a few weeks...) even has sparse matrix-multiplication (!!) units being added in.
https://www.ibm.com/software/passportadvantage/pvu_licensing...
http://www.oracle.com/us/corporate/contracts/processor-core-...
So what would you rather buy? A 24-core SMT4 Nimbus chip, or an 12-core SMT8 Cumulus chip?
The two chips have the same number of execution units. They both support 96 threads. But the 12-core SMT8 Cumulus chip will have half the license costs from Oracle.
-------
For DB2...
The SMT8 model (E950) only has a 100x multiplier for DB2, while SMT4 models have a 70x multiplier. So you're only paying 30% more per core (E950), despite getting twice the execution resources.
Even the top end SMT8 model (E980) has a 120x multiplier. So you're still saving big bucks on licensing.
What would you rather buy? A 24-core x86 or a 12-core SMT8 POWER?
That's unified L3 cache by the way, none of the "16MB/CCX with remote off-chip L3 caches being slower than DDR4 reads" that EPYC has.
Intel Xeon Platinum 8180 only has 38.5MB L3 across 28 cores / 56-threads. With 6-memory controllers at 2666 MHz (21GBps per stick), that's 126 GBps bandwidth.
AMD EPYC really only has 16MB L3 per CCX (because the other L3 cache is "remote" and slower than DDR4). With 8 memory controllers at 2666 MHz, we're at 170 GBps bandwidth on 64-threads.
If we're talking about "best CPU for in-memory database", its pretty clear that POWER9 / POWER10 is the winner. You get the fewest cores (license-cost hack) with the highest L3 and RAM bandwidths, with the most threads supported.
--------
On the other hand, x86 has far superior single-threaded performance, far superior SIMD units, and is generally cheaper. For compute-heavy situations (raytracing, H265 encoding, etc. etc.) the x86 is superior.
But as far as being a thin processor supporting as much memory bandwidth as possible (with the lowest per-core licensing costs), POWER9 / POWER10 clearly wins.
And again: those SMT8 cores are no slouch. They can handle 4-threads with very little slowdown, and the full 8-threads only has a modest slowdown. They're designed to execute many threads, instead of speeding up a single thread (which happens to be really good for databases anyway, where your CPU will spend large amounts of time waiting on RAM to respond instead of computing).
SMT8 chips are also known as "Scale up" chips. The Summit supercomputer was made with SMT4 chips by the way.
I never used an SMT8, but from the documents... its really similar to two SMT4 cores working together (at least on Power9). SMT8 chips are a completely different core than SMT4 chips, with double the execution resources, double the decoder width, double everything. SMT8 is a really, really fat core.
If a given IBM server runs PowerVM, it's SMT8. You may find this table of mine helpful (assembled from various sources and partially inferred, so accuracy not guaranteed, but it represents my understanding): https://www.devever.net/~hl/f/SERVERS
Too late for me to edit this: but POWER10 has a full 128-bit vector unit per slice now. So one more x2, for 3.8 TFlops single-precision on 30x SMT8 or 60x SMT4.
So I was off by a factor of 2x in my earlier calculation. Power10 has a dedicated matrix-multiplication unit, but I consider a full matrix-multiplication to be highly specialized (comparable to a TPU or Tensor-core), so its not really something to compare flop-per-flop except against other matrix-multiplication units.
But it seems to be made up with their implementation of SMT4. The 2nd thread on a core didn't have much slowdown at all, while thread3 and thread4 per core barely affected performance.
It seems like POWER9 at least, benefits from running significantly more threads per core (at least compared to Xeon or AMD).
EDIT: It should be noted that IBM's 128-bit vector units are downright terrible compared to Intel's 512-bit or AMD's 256-bit vector units. SIMD compute is the weakest point of the Power9, and probably the Power10. They'll be the worst at multimedia performance (or other code using SIMD units: Raytracing, graphics, etc. etc.).
Power9's best use case was highly-threaded 64-bit code without SIMD. Power10 looks like the SIMD units are improving, but they're still grossly undersized compared to AMD or Intel SIMD units.
The first, and third, threads use the "Left Superslice", while the second and fourth threads use the "Right Superslice". All four threads share a decoder (Bulldozer style).
1/4th of the branch predictor (EAT) is given to each of the 4x threads per core.
Register rename buffer is shared 2-threads at a time. (Two threads use the "left superslice", two other threads use the "right superslice"). An SMT1 mode, the single thread can use all 4 resources simultaneously.
A lot of the out-of-order stuff looks like it'd work as expected in 1-thread to 4-thread modes. At least, looking through the Power9 user guide / in theory.
--------
Honestly, I think the weirdest thing about POWER9 is the 2-cycle minimum latency (even on simple instructions like ADD and XOR). With that kind of latency, I bet that a number of inner-loops and code needs 2-threads loaded on the core, just to stay fully fed.
That'd be my theory for why 2-threads seem to be needed before POWER9 cores feel like they're being utilized well.
Obviously, POWER10 probably will change some of these details. But I'd expect POWER10 to largely be the same as POWER9 (aside from being bigger, faster, more efficient).
Edit: Turns out there is a major section on it below.
I would happily experiment with one of these at our HPC cluster(a small group at an University), but the idea of talking to a sales rep to even figure out what it would cost puts me off completely, ignoring the licenses for most interesting things to do with the hardware you buy.
I wish t. power boxes were as easy to buy as x86 boxes. Simply configure, get an idea of a price, talk to the distributor and place an order.
You might also not need the Radeon WX7100 Pro, if a consumer level graphic card is enough there are sevaral options: https://wiki.raptorcs.com/wiki/POWER9_Hardware_Compatibility...
The more you’re ready to do-it-yourself, the more you can save. You absolutely need the motherboard and IBM chip, and then you can get other things like memory, storage, graphic card, etc. used for much cheaper. Only the memory is a bit special as it has to be registered, but because of that you might be able to find good deals since it can’t be used in standard desktop computers.
You can get more performance for the same price with AMD/Intel, but for those interested this platform is not that out of reach. I use it as my workstation and I’m happy with it.
Not unreasonable. $500 for the 4-core processor or $800 for the 8 core. Basically what you would have paid for a decent system decades ago, without adjusting for inflation.
That’s 8 cores, which since Power9 has SMT4 means 32 threads. The 4 core CPU bundle is a bit cheaper but when you add the 2u CPU cooler its price gets very close, so it’s not worth it unless your computer case is too narrow for the 3u (≃15 cm clearance, https://en.wikipedia.org/wiki/Rack_unit) cooler included with the 8 core CPU.
It won’t run AmigaOS unless on an emulator[1], and then you’re better off with AMD/Intel: you would lose more performance in emulating the CPU, but you would get more raw performance for your money anyway.
[1] https://forums.raptorcs.com/index.php/topic,75.0.html
If you want to run AmigaOS software natively on PowerPC your best option is a PPC Mac running MorphOS. It even runs on the PowerBook G4: https://ddg.co/?q=morphos+powerbook+g4
Or if a well-integrated emulator suits you better, take a look at AmiKit: https://www.amikit.amiga.sk/
I don’t think that this CPU would be a net advantage for those use cases. I did use Amiga computers even as Commodore went bust, hanging on, but the approach I’m going for now is FPGA implementations: https://misterfpga.org/ http://www.apollo-core.com/index.htm?page=preorder
I’m in the waiting list for Vampire4 Standalone, but for now I’ve got a MiSTeR.
So no, don’t buy a Power9 hoping to run AmigaOS on it any better than on a mainstream PC.
P.S.: the $1310.99 price mentioned by aww_dang must be from a few months ago, the current price displayed in Raptor’s website is $1732.07. Prices went up due to the current COVID-19-induced crisis, which also made Raptor cancel their upcoming Condor platform (ATX Power9 but LaGrange-based: twice the memory bandwidth): https://www.talospace.com/2020/07/condor-cancelled.html
P.S.#2: for those in Europe, Vikings (https://vikings.net/) has been trying for a while to offer them here and they might be able do so it soon: https://www.talospace.com/2020/08/vikings-upcoming-openpower...
P.S.#3: used BlackBird with 8C Power9 for sale in Germany: https://forums.raptorcs.com/index.php/topic,153.0.html
In the end, it was great but not magical, so the difficulty of acquiring offsets the benefits IMO.
I've spent some time looking to see if it would be possible to spin up a POWER based VM in the cloud (just out of curiosity really). While it seems possible in theory in the IBM Cloud, it seems IBM themselves is only interested in offering this to their enterprise customers moving to the cloud, focusing it on getting people to move IBM's AIX/iSeries lock-in to the cloud. When looking at it before, I was not able to spin up a POWER VM from a regular IBM Cloud account at least.
There might be a bit more interest in POWER if it wasn't so damn inaccessible but it really is. If the easiest way to get into POWER is paying thousands of dollars to a company retrofitting IBM's decidedly non-desktop hardware into desktop hardware (talking about Raptor CS), your architecture is doomed to wait for all your enterprise customers to move to amd64 commodity hardware. Maybe that is IBM's goal even, I don't know.
yup -- not turn key at all -- "contact sales" but it exists: https://cloud.google.com/ibm
Unfortunately completing this flow - with the goal to get a temporary shell on a Power system running some Linux distro - requires me to authorize a payment of over 1300 dollars before I even select an image. That's for "reserving" 1 POWER core and 2GB of RAM... As an individual developer playing with this, that's way over my head. I'd understand very high end cloud pricing for that (tens of dollars per hour, I'd be happy to pay that for messing around a little), but this isn't even a cloud machine pricing model. For reference, $1300 is over half the price of getting a Talos motherboard with an 8-core POWER9 CPU you then actually own.
It seems IBM does not have the infrastructure or volume to permit cloud pricing here. I understand they might not have this, but in my opinion they need to work on this to make it accessible.
If you don’t reward good ideas (they don’t even need to be particularly good or novel, just common sense), you’ll have a company trying to grow something that’s reached its peak usage with motivational speeches, very inspiring leaders and good old pressure on employees as a form of local optimization.
I just don't think most companies run this way, maybe it's the military way of thinking of strict hierarchy rather than a market where the best ideas can develop or where you can at least break out of the hierarchy to get something started.
The MBAs and bean counters would probably argue that it's important to get certain unpopular things done, but I think a market would solve that too, as soon as something becomes a bottle neck, someone would step in.
My impression is that organizations are far too centered around VPs and directors who get to carry on without having to prove themselves again in the new situations they're in.
They care about the quarterly report and right now the best way to improve short-term performance is to milk your existing locked-in customers for as much money as possible.
That's why you'll never see a Netflix or a Whatsapp on IBM infra. But banks, insurance companies, medical companies etc are still large contributors to IBMs revenue streams. If your idea of software development is agile teams and capable programmers churning out code to power your business then you're not an IBM customer or even a prospect.
If your idea of software is 5000 programmers as interchangeable cogs in a machine with an annual release and three month acceptance cycles then IBM is where you'll probably end up.
They could be using that to do what AWS does with Graviton2 - cost control their completely integrated stack and make it a competitive advantage. Sell more performance per dollar despite having a non-amd64 architecture. Use it to give everyone more choice and competition. But instead they mostly use it to lock in their old (or new?) enterprise customers. The irony is that these two models could easily coexist, but they don't seem to understand the former and understand the latter very well.
And I can't help but think the latter is a losing model, as my strong impression is that the movement in the enterprise is away from "enterprise hardware", and towards commodity hardware. IBM needs to work on becoming commodity, in my opinion.
I know some of the IBM story from very close and it takes a certain attitude to even want to work there.
If IBM had even just kept pace and stayed behind AWS with Softlayer after they acquired them, they would have a healthy cloud business by now. It might be because I maintained a system on Softlayer both pre and post acquisition - so I was really close and able to see what was happening - but they squandered a huge opportunity there.
edit: oh hey, you're ahead of me :D https://news.ycombinator.com/item?id=24185759
Well besides the fact that 1) IBM is the only party with both the capabilities and interest in making high performance Power based products [1], and 2) evidently does not understand how to (or why to) invest in bringing this to a general audience. I don't really care if they would do this in IBM Cloud or with other cloud infrastructure companies, but they don't seem to be doing either. In addition 3) why would other parties be interested in running Power when they can run amd64 or arm? It certainly doesn't look to have a price advantage...
IBM really needs to shepherd Power well - they're the only ones that can do it. But I can't help but thinking they seem to be leading it to the grave despite apparently very capable engineering.
[1]: And why would anyone but IBM go for Power at this point if they can have ARM too? An open license matters very little when compared to ARM's mindshare and momentum.
Which is a huge shame, the POWER ISA per se is mostly fine, and their commitment to open firmware etc. for trustworthy computing (see Raptor) is a niche that for some reason interests nobody else.
If they want to turn it around and compete with x86, ARM and maybe even RISC-V on the low end, they need to commit to openpower. Get some interesting cores open sourced under the openpower umbrella, docs, open source bus interfaces for connecting stuff on a SOC, etc. And get some decent priced hardware into the hands of hobbyists and as dev boards for embedded.
Besides that, some of the atrocities that I've seen IBM and their partners commit are a good warning that you want to stay far away from them. Lest you be Watsonized and made dependent on marketing-masquerading-as-technology.
https://developer.ibm.com/linuxonpower/cloud-resources/
Each of these do some filtering because they give out resources only to find people bitcoin mining. Also, a good number of experimenters give up on the first obstacle they encounter or aren't really well-versed in benchmarking / architectural differences (lots of folks running microbenchmarks). There are some incredible resources available (many free) for those looking for a partner and not just a box. Good luck in finding the right partner for your projects.
OTOH, getting into the developer program requires talking to their sales/marketing droids too. But, if your a software OEM or have a SAS product, they were (are?) hungry and actively doing their best to recruit companies to their platforms. So, you spend a couple hours answering their questions, and in return you get some pretty steep hardware/software discounts and they will loan you machines.
To give you an idea, I can log in into a local vendor's website, for example, https://www.atea.dk/eshop/products/?filters=S_sr650 and quickly get an idea of what an SR650 Lenovo rack server would cost me. My configurations would obviously change the cost, which is where I talk to the vendor to get the real costs.
I'm not saying talking to sales people is necessarily a pleasant and easy task, of course, especially as you typically need to know more about it than they do.
Why making it so difficult to navigate? Just give resources an price.
I have to invest so much effort into an infrastructure I don't even know If I'd like to have.
From memory, about 8 IBM people showed up. They didn't seem to actually know each other, but were from several different groups within IBM.
We sat down, and started by explaining what we did with our existing x86-64 systems, and what we thought we'd like to try with the POWER system. We asked for a single box to evaluate, roughly the equivalent to our existing HP DL380 dual-Xeon boxes.
The folks from IBM then spent the next 40 minutes arguing with each other about exactly which system we should be using. Five minutes before the meeting was scheduled to end, one of them took charge and said they'd figure it out offline, and get back to use with the details.
Several more rounds of email were exchanged, but we never actually got to the point of being told what system we could have, or what its specs were, let alone actually being able to physically get one and test it.
It was perhaps the most absurd situation I've seen in 30 years in the industry.
Rather more than 20 years ago, a DEC/COMPAQ salesperson cold-called me to see if the ISP I was working at would like to switch to Alpha servers. After about 20 minutes, he offered a six-month free loan of a mid-range server -- probably $10-15K, I don't recall. It arrived a week later. We determined the hardware was pretty nice but the operating system was a major PITA -- this was when you expected to compile a large fraction of your software -- so it mostly sat on someone's desk for the last four months of the loan before the salesperson came to claim it.
But, yes: if you want to market hardware that nobody else can supply, you need to get it into the hands of people who will use it and evangelize for it.
The first one was the latest server that could run Series i. No big deal, we were coming to the end of our maintenance contract and would probably be buying a new machine in the next year or so.
The second one was a 2U, 48-socket POWER server. They bragged about it running thousands of Linux VM's at a time. I found it a bit odd because nobody who would be running this particular ERP software would be running thousands of Linux VM's.
[edit]We bought our new iSeries (Power9-based) through a VAR this year with an IBM rep helping us with what we actually needed. It was a bit of a drawn-out experience, but overall, I wasn't displeased. It was easier than dealing with some PC vendors (looking at you Dell and HP). I would imagine it will be another 15 years before we have to buy another one.[/edit]
When I worked at IBM, I usually called in a favor from friends at a local VAR whenever I needed to order something to get a BoM, because even internally the process was opaque.
https://osuosl.org/services/powerdev/request_hosting/
There are also these https://raptorcs.com/TALOSII/
You can also run under qemu, but I don't know how solid that is, and there is at least one simulator.
(Disclaimer: I work for SUSE.)
When your are already invested it is far safer to stay with what works for you. My downtime is measured in hours per year and all of it has been scheduled across the last seven years; that was the last time we had an unscheduled outage. You can achieve this will all types of hardware; well maybe not all types; but it is easier with some than compared to others and expectations are certainly much higher.
We joke at work that we get forgotten all the time because our daily ready for business meetings; all groups reporting in; never see us mentioned except to state plans for upcoming quarterly maintenance. A pleasant state to be in
This is probably a bit picky since you don't mention what those machines do, but that's not even "four nines" in the high-availability scale. I'm not sure that's something to boast about for a Fortune 500 company or any kind of praise for the IBM series. Maybe you meant hours over the last seven years?
If the "hours per year" is true, they're actually achieving 100% uptime.
Fucking big deal, if you ask me.
We ended up using ActiveMQ, and literally that evening I had things talking to each other. No contracts, no hassle, just downloaded and got to work.
A bit more on-topic in regards to IBM though- we buy licenses from them directly, and the past few months they have sent kind of bizarre emails saying I could save money if I bought them through a third party partner. I just forwarded the mail to the finance guys, but I can't imagine why IBM would introduce a middle man into our relationship with them and how that additional layer in between could end up saving us money.
Wow, I didn't realize their culture had changed that much. They were a sales oriented company since the days of Thomas Watson Sr.
[1] Yes I know IBM used to sell everything down to the pens and pencils, but virtually always to support some very expensive other purchase. [2] Yes I know IBM has things called "SMB sales" or some variation. From my experience at IBM, they were either targeting some specific product/market combo, or they were a bad joke; not exactly the A-team. YMMV.
Not for a while: https://twitter.com/RaptorCompSys/status/1295364416469377026
The systems are still quite expensive, but likely within the budget of a University research lab.
I've recently tried to engage with EventMobi, a company that supports virtual events.
I'm ready to spend thousands if necessary.
However:
#1: There is no way to just sign up, which is off-putting.
#2: Requesting a demo just put me in their funnel with canned email messages. And after over a week, I have yet to be contacted by a live human, despite sending emails.
I feel companies leave a lot of money on the table by not having some sort of self-driven onboarding.
Sales reps are humans. They get busy. Forget to call back, etc.
I feel like there should ALWAYS be at least some sort of self-driven flow at the low end. Even if sales are required at the high-end.
Otherwise, it seems, money is always being left on the table.
The dual chip module has 30 SMT8 cores running at 3+GHz, capable of 64 FP64 FLOPS/cycle when using the matrix unit. That gives 5.7TF of peak FP64 performance (Compared to 19.5TF on NVIDIA A100 when using tensor cores, and 9.7TF on A100 when not using tensor cores).
They say it has 3x the "general purpose socket performance" of power9 in FP workloads. Trying to make sense of this from the other data, they have 15 SMT8 cores per chip (12 on Power9). The single chip module runs at "4+"GHz and dual chip at "3+"Ghz. (4GHz on Power9). Each SMT8 core has 30% additional performance compared to Power9 (slide 13). If I assume the lowest possible clock that gets me to 2.4x comparing the dual chip modules to the previous single chip modules, whereas assuming 3.75GHz clocks would give 3x.
It's good to see that POWER still excels in important use cases.
The native interconnect in 10 looks interesting.
Cheaper $ cost per TFLOPs to make up for the trouble of dealing with a specialty instruction set? Speed of certain specialized computations that cannot be matched by alternatives?
Or how would one summarize it?
And yes, when used ‘correctly’, these systems can be very fast... in the steady marathon kind of way rather than the spasmodic sprint-racer clock-boosting-and-throttling manner of today’s mainline chips.
How do you reconcile this comment with the one from reacharavindh?
What use is the "Open Architecture" part if they're super expensive (ok, maybe you can ignore this part) and you have to go through sales representatives for a simple sale?
Those are still barriers to entry, even if they're not technical.
1. Is any of these "Open" architectures actually used in production anywhere serious when not implemented by their creators? I'm actually interested to know of examples.
2. How do we know that the actual chip IBM provides is the thing in the spec? The comparison was with Intel, how can we prove that there are no backdoors for PowerPC? If we can't prove, does it matter if it's "Open"?
I never claimed it was, and this is starting to reek of a straw-man argument where you’re opposing a statement I haven’t actually made.
(1) Yes, I am aware of situations where this architecture has been chosen by a body that isn’t a chief implementor, and no, I am not at liberty to discuss it.
(2) Having an open spec to compare against, even though I actually don’t know how, is already another plane of existence compared to not having something to compare against.
Decapping and microscopy? Pushing edge cases onto the chip and comparing expected outputs? Implementing all or part on an FPGA and seeing how they compare at a severely clock-reduced rate? It’s well beyond my technical ability, but it’s not beyond expert technicians’ abilities. That’s the key point.
EDIT: Also you can set your own keys for the root trust, and remove others’. That’s very important, and radically orthogonal to the competition of ARM and x86₆₄.
> ‘Open’ means that you’re allowed to understand exactly how it works and that there’s no mysteries. It means having the blueprints of the machine, not a free machine.
No, this is 100% incorrect.
The Power ISA, i.e., the software/hardware interface of the CPU, is open source. This means that if you want to build a Power CPU that implements its software interface, you can do so "for free".
That's it. You don't get "the blueprints of the machine", you cannot look into how the CPU work internally and understand it, etc.
That's like having a standard API that anybody can implement, e.g., the C standard library, but which Apple, Microsoft, etc. ship as a black box binary blob, so you can't understand their implementation, search/fix bugs, etc.
So no, your claim is completely incorrect. The benefits of an open ISA only apply to those wanting to build their own CPUs, which for Power is just not even a handful of companies, none of them making their blueprints of their CPUs openly available...
For end users, your machine is as open/closed on a system with an open ISA like in one with a closed one. People paying 10k$ for a Raptor II in the name of openness are throwing their money away.
This is a completely different situation than, e.g., RISC-V, where not only the ISA is open-source, but the VHDL implementation of many RISC-V cores is also open source, and you can buy those cores today.
If your organization has an existing relationship with IBM or Red Hat, Power CPUs are part of an integrated bundle moving forward.
Which is pretty much how it works in the Summit supercomputer (POWER9).
--------
From a CPU-perspective, its going to be more costly than a Xeon or EPYC and not as fast at crunching numbers. But POWER9 (and I expect POWER10) usually had the best L3 cache and RAM performance.
The 1TB/sec OpenCAPI link to FPGAs or GPUs continues that tradition. That's an absurdly huge communication path between CPUs and/or GPUs or whatever else is on the motherboard.
Lots of transistors and opcodes have been sacrificed for fancy things like transactional memory, runtime instrumentation and other features but fundamentals haven't improved requiring expensive compiler opts which interpreters don't do and are expensive for JITs.
The Intel chips did the fundamentals better, has POWER caught up?
The "Power" business is still doing ok; but I'd bet in a few years one of the other big guys will go at it (maybe Nvidia?) and start eating at their market share.
It's one of several reasons why smaller chips are more area-efficient to make, and one of several reasons why the major semiconductor manufacturers have been so interested lately in building chips out of smaller pieces manufactured separately rather than one big monolithic chip.
For the uninitiated, yeah, some dies can die during dicing. But I think you'd have trouble finding the cracks, and then it's just infeasible to precisely cut both halves where they would need to be cut, then reattach them. The issues would be the cut thickness, not damaging the circuits near it, precisely aligning the circuits, and then electrically connecting the circuits.
Alignment is probably the hardest part, we can barely do it for flip-chip wafers/silicon interposers on the order of the µm, imagine doing it at less than 7nm, which is the transistor pitch here.
I think they do sometimes put test features in the corners if there's space. The electrical properties of the die can vary in interesting ways [1], but the edges are usually worse than the center.
[1] https://www.google.com/search?tbm=isch&q=wafer+defect+patter...
CDs and DVDs write in a circular pattern starting from the middle going outwards, but the actual chips on these wafers seem to be their own individual squares.
https://www.ebay.com/sch/i.html?_nkw=wafer+chip&_trksid=p238...
Poured in Acryl some would make cool plates.
https://en.wikipedia.org/wiki/Wafer_(electronics)#Proposed_4...
From that one slices the wafers and then the processors get made. The ”extra” ones in the edges have pretty much zero marginal cost.
https://www.youtube.com/watch?v=8QKzS_w_Ko0
Silicon Ribbons begin at 10:10.
https://en.wikipedia.org/wiki/String_ribbon
>Ribbon solar cells are a 1970s technology most recently sold by Evergreen Solar (which is now in receivership, i.e. bankrupt and liquidated), among other manufacturers.
https://en.wikipedia.org/wiki/Crystalline_silicon#PV_industr...
>ribbon silicon (ribbon-Si), has currently no market
The wafer is round because it's cut from a cylinder of silicon. And the cylinder is a cylinder because spinning is involved in the process to make it. Hence, thanks to centripetal force, it ends up being round!
I would guess that dies are built from modular sections (e.g. SRAM cells), and it’s important that two identical modules perform identically - signal propagation time is relevant at this scale, so the shape and layout of each module must be identical. I would further guess that rectangular layouts are easiest to reason about, easiest to make masks for, easiest to pack efficiently at the transistor level, and easiest to test.
But I don’t know of a fundamental reason why a sufficiently advanced VHDL “compiler” couldn’t produce hex-cell or even circular layouts.
But - as you say - the modular sections are rectangular, and for most applications there's no good reason to make the dies any other shape.
There's actually a patent for hex-cell chips, but it doesn't seem to have been used for any significant projects.
Some metallic contacts (mostly aluminium), silicon oxide and other residues are likely present as well, depending on the masking process.
[1] https://en.wikipedia.org/wiki/Doping_(semiconductor)#Silicon...
That's the other advantages of chiplet design: maximize yield (a small defect renders a much smaller chip unusable), and much more granular binning (easier to sort out good/worse chips, due to placement and random issues during fabrication). Not to mention you have a much more modular design at the end, where you only have to change the cheaper (not 7nm) silicon interposer.
BTW the dicing used (in the 1970s) to be done partly by hand. You can see a video of someone doing it here: https://youtu.be/HW5Fvk8FNOQ?t=978
IBM Power CPUs are now 2nd and 3rd.
Also, SPARC is spelled SPARC.
IBM transferred its own chip fab business to Global Foundries several years ago and it was my understanding that they were tied to them for the following 10 years. But Global Foundries announced they were abandoning EUV so I don't think they're going to be producing 7nm chips.
> Samsung Electronics will manufacture the IBM POWER10 processor, combining Samsung's industry-leading semiconductor manufacturing technology with IBM's CPU designs.
I wonder how they got out of their deal with Global Foundries.
I wonder because NVIDIA, AMD and others requiring that process all seem to land at TSMC. Qualcomm, too, but that's hardly surprising.
Apple moved to TSMC in Taiwan not too long after Tim Cook appeared on CBS claiming that the engines of their mobile devices were made in US, almost 6-7 years ago. Apple's share of Samsung's production isn't probably much these days. But they are still #2 behind TSMC and Samsung also announced recently that they are investing $100B for next 10 years in logic business which includes their foundry.
But A8 is made by TSMC, A9 has two versions, APL0898 by Samsung, APL1022 by TSMC. There were some debates on which one is better.
After that, all Ax process are made by TSMC.
Does this work with process isolation? I.E. can I make it so that each process's memory is encrypted with a different key, to prevent snooping by other processes? How (if at all) does that work with debuggers?
It's typically implemented as an extension of the virtual memory page table, and conceptually it wouldn't be too difficult to have finer-grained keys, such as one for the kernel and one for user mode processes, or even one per process.
https://jeffhendricks.net/wp-content/uploads/2019/04/Powerma...
If they deliver a significant increase in performance - even if it's only for a few specific use cases - the ripples will be interesting to watch play out for decades to come.
Can someone say what co-optimized mean here? Is this just bad marketing speak? What is intended to mean if so?