IBM mainframe beats other platforms for cost and performance
planetmainframe.com
planetmainframe.com
I also struggle to believe that oversubscribing the z CPUs 66:1 isn't going to result in performance issues. The documents on limits aren't clear but I foresee that mainframe spending a LOT of time on locks even with all the work IBM has done to help manage them.
This whole thing just really smells like a bad advertisement for IBM, which aligns with them announcing their latest gen z processor last week.
Yes, the kernels need to be "snowflakes", but that's not necessarily an issue if this is functionality you really need.
While an LPAR could technically be HVM, the nature of the changes required to an OS make it PV-like.
We just paid just under $50k per server for on-prem physical servers that have 144 cores per server and 1.5 TB of RAM each.
For 12 of them it would be less than $600k and you'd get a total of 1,728 cores and 18 TB of RAM to compare against the mainframe's 30 cores and 2 TB of RAM for over 4x the cost!
Apple’s Macintosh made waves with its cheeky hello world demo: “Never trust a computer you can’t lift” (…and defenestrate).
Similarly, with an IBM big iron box, if something goes wrong, I sure hope the 6-figure support contract gets it fixed - probably within a few hours if you’re in Manhattan or the SF Bay, but what about elsewhere? Whereas with commodity HP/Dell/Supermicro pizzaboxes in a HA configuration I don’t need to worry about needing a same-day fix for hardware issues - I could put things off for months, even - and probably do it myself and spend the support contract money on a Plaid Model S and have money left-over for a Mac Pro with XDR display (with the cool stand!).
I know IBM’s srsbsns machines have their own redundancies - including being able to hot-swap literally everything, it still feels like open alternatives are overall better in the long-run.
That aside, samepage sharing has been a thing in virtualization for more than a decade. If you have 800 VMs running RHEL8 for POWER, a significant amount of the memory load is kept in identical pages, which lowers the burden of oversubscription. The recommended workload is unlikely to be 1000 unicorn/pet VMs, and more likely to be a large number of VMs running the same base OS, and some of the same workloads.
Synchronous interrupt locking is also a solved problem.
0: https://www.kernel.org/doc/html/latest/admin-guide/mm/ksm.ht...
The CPU cycles are a non-issue for the configuration on this mainframe, and it's fair to assume that, since IBM developed the silicon and IBM engineers wrote the hypervisor layer, that most of the detrimental effects of "we need this to work across multiple CPU vendors and multiple generations of the architecture" are covered.
Even x86 CPUs get improvements to cache coherency in memory dedupe scenarios with generational improvement, and intelligent NUMA topology layout helps a lot. It's also worth noting that it's a mainframe, so essentially every configuration has a large number of other processors with a couple of instructions disabled so they don't "count" as CPUs for licensing, but they still perform work. I would not be surprised in the least if there were 1+ co-processors dedicated to offloading operations exactly like "LPAR || z/VM memory dedupe"
http://www.redbooks.ibm.com/redpapers/pdfs/redp4827.pdf
[old]
It is called active memory sharing in the POWER world
https://www.ibm.com/docs/en/power9/9080-M9S?topic=sharing-ma...
Slide 17:
http://www.vm.ibm.com/library/presentations/syslimit.pdf
>Synchronous interrupt locking is also a solved problem.
No, it really isn't. I have not seen a production implementation of virtualization with > 4:1 CPU oversubscription that doesn't eventually or immediately have significant performance issues with database workloads.
This discussion is about the configuration of the systems, not this specific scenario. While this may have been the scenario outlined for you, it wasn't apparent from your comment.
Production workloads on busy servers have very different requirements than consolidation, VDI, resiliency/redundancy, hardware abstraction, etc. Every workload needs a different evaluation.
All of these VMs running production Ora/OAM/OEM/Websphere workloads? No, don't oversubscribe. Some of them running similar workloads? More oversubscribe is ok. Few of them? Lots of oversubscribe.
Similar for interrupt locking. The "old" synchronous interrupt lock I was speaking about was "I have a bunch of VMs with CPU oversubscribe, and even if they're doing nothing, 30% or more of your CPU time goes to interrupt scheduling so every vCPU for a given VM can schedule simultaneously". This is solved.
"I'm running a CPU-intensive workload on a massively oversubscribed server" is not, and we really shouldn't expect it to be.
I am having a generalized discussion about virtualization oversubscribe. You are having a specific discussion about CPU-heavy DB workloads. Apples cannot be compared to oranges.
>All of these VMs running production Ora/OAM/OEM/Websphere workloads? No, don't oversubscribe. Some of them running similar workloads? More oversubscribe is ok. Few of them? Lots of oversubscribe.
All of the VMs are running Oracle/Websphere/Apache, per the link. That is what the article is recommending, which is where my skepticism came and is coming from.
>I am having a generalized discussion about virtualization oversubscribe. You are having a specific discussion about CPU-heavy DB workloads. Apples cannot be compared to oranges.
Everyone in this thread is talking about the article, I guess everyone but you.
That aside, comments on articles are sometimes more general. "How do mainframes compare in 2021?" is a conversation which is not intrinsically linked to "should you oversubscribe VMs 66:1 for this specific workload?", and I interpreted your comment at the former. I wasn't talking about the article at all.
There is a giant table in the middle of the article that is impossible to miss if you read it. It literally says # of VMs. As well as calling out VMware, and z/VM as tye hypervisors in use. The oversubscription is a simple math problem, and that's assuming best case scenario.
They’ve stopped being an innovation business.
Most applications don't need gigabytes of memory though? I don't know what one uses a mainframe for but web servers, file shares, remote login software, proxies, traffic scrubbers... lots of common applications use up to a few dozen megabytes of RAM; throughput is much more important there and memory to operate on is just for some state and buffers. Even a database with millions of rows and a dozen indexed columns, the index of that (I just checked a database of my own) is in the hundreds of megabytes, far from even one gigabyte.
It seems a bit odd to assume each server always needs multiple gigabytes of RAM.
Well, yes. Oracle charges per core, and you are running Oracle EE on an awful lot more cores. But what if you are running Postgres instead?
Just don't tell 'em that you really just installed Apache and replaced the admin login page.
Guess what happened when Oracle learned we were planning to have 3 million users...
It seems that this would never work in a fast moving and fast scaling environment, and the added 'special' sauce makes it incompatible with the broader FOSS ecosystem (including when "it really is the same as Linux on x86" - which it never is).
Say you really do run the RHEL+Apache Webserver+Websphere+MQ+Oracle stack, are you really in a position to make sane choices anyway?
Better yet, if I have 10 people doing shared services on a public cloud supporting 150 developers developing applications that run on K8S that scales from 10% to 1000% on a daily basis. How would this IBM mainframe (or even the ancient vSphere 4.0 on blades) deliver any value over K8s and a cloud? Do I fire 100 developers to then hire 20 IBM auditors, 20 IBM mainframe maintainers and 10 legacy developers? Because cost-wise that would be the same, yet output-wise it wouldn't nearly deliver the value we would have had before.
And on top of that, the data tables are images? wtf?
I'm sure you can construct a scenario where a mainframe would 'win' (as if there is such a thing, only 'fit' matters), but commodity workloads on commodity resources has 'won' a decade ago. The only ones that are stuck are specialist cases or 'classic' multinationals.
> How would this IBM mainframe (or even the ancient vSphere 4.0 on blades) deliver any value over K8s and a cloud?
There value is in mitigating risk. The engineers want the new and shiny toys and for all the right reasons (development velocity, etc), but the managers fear change bc change == risk.
> Because cost-wise that would be the same, yet output-wise it wouldn't nearly deliver the value we would have had before.
Likely, you are right, but the budget committee will exercise some financial gymnastics in order to justify the outcome they want, instead of the budget dictating decision making.
> The only ones that are stuck are specialist cases or 'classic' multinationals.
Wrt banks, they are stuck on cultural inertia and risk mitigation.
Yep. My organization is trying to move away from mainframe and COBOL due to limitations in available personnel. The systems perform well, but that doesn't matter if we can't find people to keep them running.
A fast moving environment by definition would never even consider mainframe computing. And most places who would consider mainframes are already at whatever scale they will be at the conceivable future.
I used to work for a public healthcare company in the US that used mainframes extensively and when I was there we legally couldn't open any new locations. The only way to grow the business at all was to acquire existing ones, a multi-year process. So any changes or integrations typically had 12-18 months of lead time, and this was after everything was negotiated and signed. We typically heard about it a few months before that so sometimes up to two years.
And as you can imagine with the amount of red tape and regulation around healthcare, and the requirements around being on the stock exchange, change control was a pretty arduous process.
If the VM runs Linux, it's just Linux running on s390x. It's no weirder than ARM and has been on the Linux server market for far longer. I think you can install Linux directly on an LPAR. The part admins will need to learn is zVM, which isn't much more complicated than KVM or vSphere (and is much easier than K8s). The hardest part is the jargon and acronyms: IBM invented a lot of things that other manufacturers named differently.
> and the added 'special' sauce makes it incompatible with the broader FOSS ecosystem (including when "it really is the same as Linux on x86" - which it never is)
A mainframe usually can run Linux on zVM and it behaves just like a 5.2GHz VM with a very fast IO subsystem. From the top, it just looks like a very large computer that can run a lot of VMs in a single system. IBM started hitting diminishing returns with the number of cores and that's why core counts aren't increasing as they do on ARM and x86. Under Linux, the cores do SMT2.
> How would this IBM mainframe (or even the ancient vSphere 4.0 on blades) deliver any value over K8s and a cloud?
zVM is very mature technology and Z hardware and software have been coevolving for longer than most of us have been alive. More interestingly, you can partition your mainframe (into LPARs) and have several completely separated and isolated systems. It’s below zVM, so it doesn’t even suspect it doesn’t have the machine for itself. It's very common to run production, staging and development on the same system this way, as if you had three separate machines.
> 10 legacy developers?
Unless you plan to deploy your web app written in COBOL running on CICS and zOS, I'd suggest using different tools. If you already have systems running on zOS, you can easily access them from Linux VMs hosted on the same machine, through an imaginary network that's really fast (because it's not really a network).
> Because cost-wise that would be the same, yet output-wise it wouldn't nearly deliver the value we would have had before.
As the article points out, the mainframe itself is a little more costly, but the operational cost is lower. The convenience of having a single very reliable system instead of a cluster of less reliable ones can't be ignored. As Seymour Cray once said, it's better to plow a field with two strong oxen than with 1024 chickens.
There are downsides - IBM doesn't have these machines in stock the same way you can order a Dell, and ordering one is not a simple process - they'll build one for you, tailored to your needs. They'll probably deliver it with a couple extra CPUs so that when you need to upgrade it, you just pay the license and activate those resources.
One cool thing the z15 does is to activate all processors on boot to speed up the process. After the machine is running, the parts you didn't pay for shut down and the machine continues working.
Running a mainframe comes with the challenges of finding specialists, relying on a single vendor, having to buy at least two in two different data centers for disaster recovery,etc...But purely from performance point of view, as in throughput, and also quality of service, the case is clear for the mainframe.
Spend some time reading the Z15 technical guide, and marvel at what is today the state of the art in computing. We are talking about sustained, full on 5.2 GHz, all the time. Massive IO bandwith, ECC memory everywhere, massive caches, native crypto co-processors.
All this, while your x86, most likely, spends most of its time waiting on memory...
"A Crash Course in Modern Hardware by Cliff Click"
https://www.youtube.com/watch?v=OFgxAFdxYAQ
IBM z15 (8561) Technical Guide
https://www.redbooks.ibm.com/abstracts/sg248851.html?Open=
https://www.redbooks.ibm.com/redbooks/pdfs/sg248851.pdf
What cloud vendors are doing is using, sometimes, custom hardware to create a mainframe at cloud scale. Its just due to the sheer incompetence of IBM management, that they were unable to leverage their mainframe technologies, and offer cloud offerings that would be competitive on a price per compute or storage unit.
But they could have created a computing offer, using the pure core performance of the Z processors and the technologies that enable things like z/OS Sysplex and do it at cloud scale. They could then maybe be competitive on a price scale:
https://www.ibm.com/docs/en/zos-basic-skills?topic=sysplex-z...
This would be similar to AWS telling you could have your CPU core stretched across two availability zones.
No, it isn't clear. Show the benchmarks.
We can start here:
"IBM z15 Performance of Cryptographic Operations"
https://www.ibm.com/downloads/cas/6K2653EJ
Can your x86 do this?
Edit: "Introducing the new IBM z15 T02"
https://www.ibm.com/training/pdfs/The_New_IBM_z15_A-technica...
==========
5.2 GHz core frequency
• On-chip compression accelerator (NXU)
• On Core L1/L2 Cache
• L2-I from 2MB to 4MB per core
• On chip L3 Cache
• L3 from 128MB to 256MB per chip
960 MB shared eDRAM L4 Cache
=========
https://github.com/IBM/IBM-Z-zOS/blob/main/zOS-Education/Hin...
https://community.ibm.com/HigherLogic/System/DownloadDocumen...
So, I am guessing actual head to head benchmarks must be quite bad.
Your latest Intel® Xeon® E-2386G Processor, not available yet...has a base frequency of 3.50 GHz and a Max Turbo Frequency 5.10 GHz
So even with the latest...your Turbo frequency... is the base from the mainframe core...
And if you happen to be running a memory bound benchmark, with the massive caches in the mainframe, you will be crushed.
I can easily write a memory bound benchmark that makes all CPU cache irrelevant. The only thing that actually matters is real world performance which is why actual benchmarks are so critical.
> Xeon® E-2386G Processor
That’s a chip from 2018. Anyway, you can talk about CPU clock speeds all day, but it doesn’t mean much between different CPU families.
Wrong choice indeed the Xeon® E-2386G Processor its not what I was looking for. Instead of going for the latest lets review the latest offers from Intel:
https://en.wikipedia.org/wiki/List_of_Intel_processors
A few models that have a peak turbo frequency approaching the sustained frequency of Z core. Also, of course due to different architectures, internal caches, we should not be comparing purely on frequency. Although the mainframe is again wining here in frequency and cache sizes...
Also agree that what counts are real workloads with real commercial applications. I have so far posted at least three but here is another one:
"IBM MQ for z/OS on z15 Performance"
They won't be mainframes, because as the author does correctly point out, mainframes are unmatched at large-scale, high throughput, resilient, stateful, transaction-heavy workloads. While that is part of what the major cloud providers support, it is not everything. Because it is not everything, the systems the major cloud providers design and build will look different from mainframes, but I think they will have key similarities: virtualization and security features provided by the hardware; high reliability through redundancy and in-hardware error-checking; easy physical maintenance through hot-swappable components.
I don't think hyperscalers will ever care about this. They have millions of servers. They can just take the whole server out of commission as soon as it fails a health check. This is a critical difference between the mainframe model and the modern cloud model.
So like DynamoDB, Document DB, SQS, Service bus, and vast majority of managed cloud services?
Thats literally the entire reason to use cloud - to offload management of statefull services.
If I wanted just to run stateless docker containers, I can set that up myself on bare metal in a day.
Does VISA do more transactions than all of AWS combined? Is a single mainframe more reliable than the entire datacenter running AWS service? Is Bank of America's data literally a single SQL database?
Commodity hardware goes into their racks, open firmware and bootloader, etc. etc. and their podcast is a must-listen for any computing enthusiast.
These numbers don't make sense to me.
IBM is the only company that can sell you a fully licensed (you need a license to activate CPUs) IBM mainframe.
And if not, can we have a law similar to the first-sale doctrine here to limit the control of the manufacturer?
Starting with making a comparison on list pricing, when commodity hardware has huge discounts, mainframes don't. Then moving on to comparisons of running Oracle on a slew of tiny VMs with ~2GB RAM each. And, no benchmarks anyway, so what does "beats" mean?
It also should not focus on HUGE cloud providers like Amazon, Google, Microsoft.
Estimating licensing, maintenance costs for running at that scale would need access to their internal documents (as far as I know, none of them have published all the numbers).
Hardware costs are hard to pin down. These companies have special specs for what they want. Having custom computers produced should put the price higher than off the rack, but given how mnany servers they buy that probably changess.
In today's world, cloud can run high volume transaction loads and most recent IBM z15 and IBM LinuxOne can run normal cloud tasks.
I believe that a proper comparison in 2021 would be a lot closer than what most x64 people would expect.
As long as we ignore AWS,Azure,Googe scale and think of smaller "in house" cloud like datacenters it should be part of the estimation.
In my opinion DB2 running on a z15/zOS is hard to beat on uptime, reliability, and throughput. There would also be less need for sharding and clustering headaches.
A z15 can also be loaded up with Specialty Engine Support like IBM Integrated Facility for Linux (IFL), or IBM zEnterprise® Application Assist Processor (zAAP)
Which means added a bunch of cpus that are dedicated to oen thing, in those mentioned above one is for running Linux and the second for running Java. All offloading the main system
One big problem in comparing is that the terminology and smenatics are quite different. Reading specs does not make it easy to compare 1 - 1.
I have hoped i would end up at a place that did use modern mainframes for "modern conventional" workloads. I havent had that chance yet.
I am pretty sure AWS for example doed not have major licensing payments for software. Their stack is mix of open source and self-written software and none of them require regular license payments.
So all the software costs that the study talks about would be reversed -- it is hard to compete with $0. And any extra work in system administration will be split over hundreds of thousands of servers.
But microscopic compared to having to run proprietary software on all servers.
System administration over hundreds of thousands of servers is a lot more complicated than one or a few z15 or Linux One machines.
In addition, z15 has battle-hardened reliability (far above AWS servers). All the infrastructure to feed and care for over hundreds of thousands of servers is a lot more complicated. Even if what you have to do is just shutdown servers that are not working and never turn them on again.
It is one of the biggest advantages of the z15 platform.
IBM does some crazy stuff with hardware acceleration and specialized units that take load off of the actual CPUs.
[0]: https://github.com/rust-lang/libz-sys/pull/72#issuecomment-9...
[1]: http://cbloomrants.blogspot.com/2020/09/how-oodle-kraken-and...
1) Single supplier. Have a falling out with IBM.... you're stuffed. that's a major business risk. With x86, don't like HP go to Dell etc.
2) Staffing! What's the relative availability of good S/390 admins and devs, compared to x86? IMHO one of the reasons first mainframes, then AS/400, then proprietary Unix lost out to Linux on x86 was ease of getting staff, which is fuelled by the fact you can start out with a home PC and learn the kind of skills you need.
You can’t just own one of something.
Plus, when you are at 95% capacity, your only option is to buy 100% more headroom.
I know this is not a popular view, but IBM is the only vendor other than Apple that can build an entire machine from the processor to the operating system. It could be interesting if their systems were cheaper so that x86 had some competition.
"Whether intentional or through ignorance, there is a great deal of bias against the mainframe. It’s too expensive! (It clearly is not.) It’s old and dusty! (Obviously not.) It’s hopelessly outdated!"
There might be 'some' bias against the mainframe purely from technology perspective, but I feel a large part of it is due to culture and people around mainframe. So the management uses the 'bias' as a guise to clean up the shop.
https://www.anandtech.com/show/16936/ibm-power10-coming-to-m...
Overall, I want to find a reason to get excited about this market segment because I'm a processor nerd, but the arguments are profoundly uncompelling.
You can run the latest zOS on Hercules/Hyperion, but IBM will not like it - and most certainly won't sell you a license. You won't be using a cool CPU, only a normal one pretending to be cool.
You can get a Linux VM from the LinuxONE Community Cloud. It's fun to play with - feels like a fast two-core (probably running SMT2 on a single one) with wicked fast IO but, apart from that, nothing too special. It has a lot of performance counters in the CPU too, so that's fun at least.
That's not what I observed. Maybe admins are more expensive, but I make a lot more as a Python person than I would be able to writing COBOL code.
iSeries (what they call the AS/400 these days) is not going anywhere soon. You can even get one from the IBM Cloud if you want to go cloud with yours. If you don't, you can still upgrade them. You can still run anything a first-gen CISC-based AS/400 could on the newest POWER9-based hardware (they are just pSeries with some hardware key that allows them to load iSeries OS).
They are not called mainframes though. They are midrange systems, the last remaining minicomputer breed. They are expensive, reliable, their OS is alien to most of us, but they don't get the mainframe badge.
(Yes, I know there are use cases where mainframes are the best. Zzz...)
You can get something similar (on a z15) for free from the LinuxONE Community Cloud.
1 vCPU
4 GB RAM
100 GB Storage (25 GB boot, 75 GB data)
is around $180/month in my local currencyIf what the article says is true, then for a similar workload, then IBM could instances ought to be much cheaper than the competition.
I only have experience with GCN, but to me it looks like GCN is much cheaper then IBM cloud, so something must be extremely off in the calculations done in the article.
https://www.ibm.com/cloud/hyper-protect-virtual-servers
The fact there is a market for them is a quite notable.
I find it perplexing too, but I guess they don't want IBM Cloud to eat away from their high-margin SKUs.
1Virtual server instance $0.106/hr 2 vCPUs 8 GiB RAM 4 Gbps Image provided CentOS Boot volume $0.018/hr 100 GB Virtual Private Cloud provided Network interface provided Apply a code Apply Please sign up to create or login before applying the promo code.
Subtotal
$89.35
Sustained usage discount -$7.87
Total estimated cost $81.48/mo
I really thought they were going to do fully disaggregated clusters, but it looks more like vxblock-with-remote-management.