Ok, maybe this is a silly question, but I am {...}.... Question: what is a house and why would I want one? In my mind I think of it as a big, but inelastic, tent. If that's true, why wouldn't I want to just get multiple tents instead?
A mainframe is the kind of computer you can buy and set up, turn on, and run for two decades (or more, if you like) non-stop. It doesn't need reboots, it most often doesn't even need to be powered down for hardware replacement or upgrades. It's the epitome of reliability.
Most people don't care about reliability. Heck - most people use Windows, so they can't even possibly imagine what reliability and consistency might be like. Therefore, most people can't even imagine wanting reliability because what they have seems "good enough".
People who run services which should never go down, on the other hand, love mainframes.
article about it: https://blog.ctm-it.com/it-support/blogs/matt-cannon/2013/49...
edit: If you want to be especially cynical, you can note that this wasn't discovered for a while because everyone running Windows reboots it regularly because of its notorious instability issues (the bug was found in 2013 and dates back to at least Vista, which was last updated in 2009).
I reboot monthly because getting owned by 0 days sucks, and at this point in life, every platform (except for OpenBSD) is having new exploits found against it at a rather fair rate. Same reason my phone gets rebooted once a month.
Aside from that, I have had uptime on Windows boxes in excess of 6 months.
Laptops on the other hand, those are more problematic, but it isn't Windows' fault if a WiFi card hard locks itself, has a dysfunctional watchdog timer on board, and needs to be power cycled to get it working again.
OSes that can hot swap or live patch kernels can brag about uptimes. The rest of us should stay humble. (And AFAIK no modern consumer based OSes distributions do that)
Windows laptops.
My MacBook has been up 56 days 14:06 hours.
Apple laptops have their fair share of hardware issues. Windows laptops tend to have more due to a shorter product cycle. Apple has longer relationships with their parts vendors and thus they have more time and leverage to force good firmware onto the multitudes of independent chips that make up a modern day computer.
Windows used to suck, but it's somewhat unfair to judge the platform of 2018 because Vista of 2006 was bad for you.
Hopefully a mainframe expert can chime in as to whether the major mainframe OSes have included similar or worse bugs.
Windows has improved since then, but it's still a very brave person that would have a single business-critical Windows machine without a few scheduled maintenance and reboot periods every year.
It is, of course, only a proxy for the real reason, which is the engineering culture that led to the bug and/or reliability. If there's better evidence that Microsoft has drastically changed its engineering culture in favor of reliability and/or IBM has done the converse [1], it could easily trump this proxy evidence.
[1] for example, if they've included the OS developers in the layoffs they've been in the news for
"Well, you want redundancy right? Well you're supposed to be redundant across AZs, and then regions, and then you're going to have to have disparate vendors to mitigate sole vendor risks." And then we're right back to hosting our own mainframes in our datacenters.
Are they, though? Commodity hardware certainly is, but it's not as if cloud providers are charging a small margin on top of that. They're charging a multiple, potentially as large as 10x.
Combined with the parent's proposed need of multi-cloud, that could turn what might otherwise be a few hundred $k of commodity servers into a few $M of cloud costs, which I understand is the OOM the cost of a mainframe.
With IaaS typical reliability issue is that whole location/datacenter/AZ goes down, with your own infrastructure the typical issue is that the colo-facility/datacenter goes down, which would be essentially identical save the fact that with IaaS there is significantly larger probability that the reason for going down is some byzantine failure of orchestration automation, which in the self-hosted case either isn't there or is under your control.
One fact of running your own infrastructure is that you should plan for hardware failures, but not stress about it too much, because even entry-level enterprise-grade hardware just does not break (and if it does you will get signs that it is going to break well in advance).
That said..
> Problem with IaaS clouds is that it is strictly less reliable than having your own infrastructure.
Although I'm a fan of running ones own hardware, I'm not sure I could make this claim. However, since I'd like to, do you have public data to back it up?
> orchestration automation, which in the self-hosted case either isn't there or is under your control.
I'm not sure how that would be different. Although a provider like AWS has portions of the automation toolchain not under your control, it's not obvious they're any more likely to fail (even due to some bizarre interoperability bug) than, say, BMC firmware, which is also not under your (full) control.
> One fact of running your own infrastructure is that you should plan for hardware failures, but not stress about it too much, because even entry-level enterprise-grade hardware just does not break (and if it does you will get signs that it is going to break well in advance).
This is something that I routinely have to point out to cloud proponents when they complain about having to "worry" about hardware failing: modern, commodity server hardware just doesn't fail often enough for it to be a significant consideration. Usually, it's just selection bias in that they remember every "nightmare" scenario from their past where hardware failed (possibly even as long as 20 years ago) but don't account for the overwhelming majority of times when it didn't.
Of course, there are notable exceptions, such as high-density "blade" or half-U servers, which often suffer from thermal design failures, but I argue that those are a departure from commodity, even if they appear identical if one squints.
Most importantly, though, it's not as if an IaaS cloud provider can somehow magically shield you from the consequences of such a failure: your VM will still go down. Sure, they have an arbitrarily large supply of spares to replace it, but you only ever need exactly 1 of those spares, and N+1 redundancy when self-hosted is very easy, if implemented merely as warm spares.
That said I believe that when your workload really necessitates such a big systems and if you can use all of the mainframes capability and have the capability to manage the mainframe (which requires Ops team with totally different skillset and mainly willingness to do such a thing), I believe that cost of mainframe will be comparable to IaaS, with self-hosted commodity HW being somewhat cheaper.
It is anecdotal, but when I used to work in more of an Ops role I cannot remember single time when server hardware failed in production without external environmental cause (there were flooded servers and servers that were DoA from manufacturer), it is somewhat surprising that this experience even extends to spinning rust harddrives, where most common causes of failure I've seen were flaky SATA/SAS connetors, followed by simply bad series (eg. Constellation ES.2) and then by extreme overheating.
> cost of mainframe will be comparable to IaaS
I'd be very interested in seeing even a rough cost comparison, since I have no experience with mainframes.
I also have, essentially, no experience with that first "if", which I'd say is a big one. Workloads that are (already) suited to that particular system design may be rare (and obviously getting rarer).
> self-hosted commodity HW being somewhat cheaper.
I'm pretty sure "somewhat" grossly understates it.
However, I realized that one of the problems here is that we're talking about these costs as if they're single numbers, rather than ranges.
For mainframes, it may as well be a single number, because there's only one vendor (for the latest hardware).
For IaaS and self-hosted, the ranges can be very broad, because it's very easy to pay a multiple of the minimum cost with merely a naive implementation. Trivial examples would be not leveraging "reserved instances" on AWS or not getting competitive quotes for self-hosted. In fact, if one removes the "commodity" constraint from self-hosted and allows "enterprise" hardware (especially storage), the top end of the range can easily balloon above the top of the IaaS range.
I've been assuming a comparison of the bottom ends of the ranges, but including the cost of the expert labor for each. What's difficult to know, of course, is how scarce the experts (capable of keeping costs near the bottom of the range) in each category actually are.
> I cannot remember single time when server hardware failed in production without external environmental cause
I think that's just a matter of too small a sample size.
> servers that were DoA from manufacturer
I wouldn't even count that, since it's not in production yet.
> it is somewhat surprising that this experience even extends to spinning rust harddrives
That's definitely too small a sample size, then. If you're not seeing at least a 1% AFR (realistically closer to 3%), you don't have enough of them or haven't been running them long enough yet.
> simply bad series (eg. Constellation ES.2)
That's not an external environmental cause, though, and counts the same as any other failure due to (presumably) a manufacturing defect (defined broadly), including RAM bit errors. It's merely something that can be engineered around with best practices.
None of this is to say that any of these inevitable failures are actually frequent or voluminous enough (even on, e.g., 5 year old hardware, which is ancient by most standards) to require outsized worry or effort/cost to mitigate/repair them.
That is certainly possible for an organization to achieve, but it isn't easy and it isn't cheap. It certainly can't be taken for granted.
I'm not sure that the competence required is organizational (e.g. people managing) so much as operational (e.g. best practices) and even technical, at least at sub-FAANMG scale.
> it isn't easy and it isn't cheap. It certainly can't be taken for granted.
I agree that it can't be taken for granted, but I disagree with it not being easy and cheap. Rather, combining the two, I don't believe it's necessarily hard nor necessarily expensive.
It just requires finding someone who both still has the competence, is willing to use it, and is willing to train others. Running your own infrastructure isn't actually difficult or complicated, but it's certainly not "sexy" and can be a bit tedious at times. That means it's possible to hire inexpensive, less (overall) experienced staff and have them handle that portion. Unfortunately, the "unsexy" part means finding someone to do the training, as well as the actual work when necessary, can be challenging, even though we're out there.
Even then, that's only necessary at substantial scale. In <1000 server environments, I've never had the hardware-specific [1] part of the work take up more than a quarter of one senior FTE (usually me).
What can get astronomically expensive is outsourcing the wrong things, though that ends up being a form of not actually running your own infrastructure (yourself).
Anecdote: I recently had a phone interview with a startup that moved from "hardware" to the cloud and the main reason cited was the inability to ramp capacity up fast enough (nor predictably fast enough), which seemed odd to me. One example of unpredictability of lead times involved a new server underperforming due to mis-applied thermal compound between the CPU and cooler, which I have never experienced [2]. I didn't ask the rhetorical question, "how could you have picked such a horrible VAR?!" Carefully re-reading the blog post about their transition gave me my "aha" moment: even though it's a company in the SFBA, their datacenter was out of state (maybe not even in a tech hub city, but it didn't specify). They were outsourcing the actual installation, running, and maintenance of their hardware to someone else, far away.
[1] for lack of a better term.. i.e. anything that an IaaS cloud provider would eliminate, including purchasing and vendor negotiations, colo space, network hardware and providers, hardware monitoring, and data destruction
[2] well, OK, that's a lie, since I've experienced it when I've personally done CPU moves/swaps/upgrades in exceptional circumstances, when I was out of practice, but I knew to test my work and caught the problem immediately. I've never had it happen with professionally-assembled systems, presumably because CPU coolers tend to arrive with the thermal compound pre-applied.
So having interconnected commodity servers will always be marginally more expensive to run than mainframe boxes of the size of fridge, which have dedicated hardware for interconnect of internal components (for example on-book CPU caches and shared RAM).
You could almost think of a mainframe as a cloud in a box, and if one isn’t realiable enough, you can always run two or more.
side note; I HATE the trend in the level of abstraction away from hardware. Seems about as short sighted as GP doctors abstracting away human anatomy. They mostly treat illnesses and proscribe medication; why should they need to know the names of bones in a human body? There are orthopedic specialists for that!
And for the elasticity part, mainframes are where the whole virtualization and hypervisors started. Although, or maybe because, purchasing and operating an mainframe is significant investment in terms of both capex and opex (it is somewhat telling that IBM's specification sheets for mainframes read somewhat like flyers for new cars, about half of the thing is about available financing methods :)).
On the other hand, you can get rack-mount x86 server with several TB of RAM for fraction of price of mainframe which is exactly the hardware for new applications that would otherwise be best served by mainframe.
I expect the largest one of these is still a fraction the size of the largest mainframe, at least for processor power (e.g. 224 cores[0] compared to 1700[1]). Maximum memory, though, is 32TB, which isn't exactly huge compared to 12TB [0] or 24TB [2].
[0] https://www.supermicro.com/products/system/7U/7089/SYS-7089P...
[1] https://en.wikipedia.org/wiki/IBM_zEnterprise_System#z14 (Not saying they're necessarily equivalent in performance, but they might be, as you point out)
[2] https://www.supermicro.com/products/system/7U/7088/SYS-7088B...
That's my general impression, as well, but I have little enough (or inadequately broad) direct experience for that impression to be a strong one.
I've also seen, second-hand (i.e. benchmarks, so not quite real-world) that cache and memory latencies (including inter-CPU/NUMA) can significantly affect OLTP performance.
If that weren't true in practice, then the sheer bandwidth one can deliver from enough SSDs over PCIe rivaling a CPU's memory bandwidth would mean main memory size is much less relevant.
The IBM mainframe system is called "Z" for zero downtime.
Banks and Airlines use systems that haven't needed to be rebooted in decades.
(This is a little oversold, but that's the selling point.)
When I worked on "big iron" (Not sure if the old hp superdomes counts mainframes..), the solution proposed for redundancy was to have two and be able to switch from one to the other..
Our development machines had to be rebooted from time to time as we turned off interrupts for some processes (quasi-realtime) and when under development some bugs could cause the processes to fail and be uninterruptible. Things stabilized quickly though.
(edit: apparently there is 4 nines, 6 nines....) https://en.wikipedia.org/wiki/High_availability
Just the other day I told the post office attendant their system was back online, as the cursor started blinking again. Telltale sign of mainframe style computing.
Mainframe is a really big server, and it is actually plenty elastic. You can divide it into smaller logical servers (LPARs).
The only downside is price. But I think for larger installations and certain workloads it will come out similar with a server fleet.
You own one because your company or agency bought one in 1970, and it is cheaper to run it than to redo everything. You may need one for certain new workloads if you’re a defense contractor.
As an employee, you’re a pure operational cost center. You have skills that few employers care about, and vendors/contractors are churning out cheap replacements for you.
You don't want one if you ask me.
I'm still not seeing what mainframes give you that a well-designed pc cluster couldn't, other than marketing BS and opportunity for grift.
Waitaminute...
[1]https://en.m.wikipedia.org/wiki/IBM_hexadecimal_floating_poi...
Regardless, I wasn't able to find in the WP entry adequate support of your assertion. (Oddly, the "[..]" were in the original WP text):
> As IBM is the only remaining provider of hardware (and only in their mainframes) using their non-standard floating-point format, no popular file format requires it; Except the FDA requires the SAS file format and "All floating-point numbers in the file are stored using the IBM mainframe representation. [..] Most platforms use the IEEE representation for floating-point numbers. [..] To assist you in reading and/or writing transport files, we are providing routines to convert from IEEE representation (either big endian or little endian) to transport representation and back again." Code for IBM's format is also available under LGPLv2.1
There's only a regulation for a file format (and a single, isolated one at that), not regulation of using/processing, and format conversion utilities are available, rendering it irrelevant.
And the other examples under Special Uses... like the GRIB data, I have toyed with reading weather model data and I just write Go code that reads the files and does conversions as necessary on x86 without having to involve a mainframe. (Although that said, I have an odd fascination with mainframes despite never using them professionally, and maintain a system using the Hercules emulator and the MVS 3.8 operating system, which was the last freely-available version of what became today's z/OS.)
- IBM legacy hexadecimal floating point you mentioned
- IEEE 754 binary floating point
- IEEE 754 decimal floating point (they were a first implementation of it)
So perhaps you were downvoted because on the mainframe, you don't have to use the legacy format at all.
Very roughly speaking, a mainframe is to a CPU what a supercomputer is to a GPU.
Sure, my big powerful desktop PC spends 90% of its life surfing the web or doing light stuff, but for that remaining 10%... if I need the power, it rises to the challenge, every time. Plus, its paid for and I don't need to answer to anybody to make use of it.
Besides the (purported) reliability reasons other commenters have brought up, because "vertical" [0] scalability is, often enough, more efficient in terms of money and (human/engineering) time.
This can hold true even if paying a premium for both the hardware and the humans (which, presumably, are more expensive due to their scarcity, as implied by the article).
The promise of "horizontal" scalability is being able to achieve arbitrary performance just by adding more units (e.g. servers), but not at any particular efficiency (slope or shape of the curve). Even that promise is suspect in the face of Amdahl's Law [1] and USL [2]. In reality, the Fallacies [3] matter, so the most general-purpose supercomputers are still bespoke and expensive, even if they're assembled from (a selection of) commodity parts.
Even just in the realm of commodity hardware, I/O-heavy loads are going to benefit from keeping everything as close together as possible, on as fat pipes as possible. Distributed databases or any distributed data handling (that doesn't involve heavy processor use, such as compression or transcoding) is what I've personally seen suffer from this issue. 20-60 "inexpensive" ($3k-4k) servers instead of 2-3 "expensive" ($20k-$60k) ones.
Cloud infrastructure further confounds any calculation, since, with the most popular provider, AWS, you'd pay 2x-10x what you would buying and operating your own hardware. That, multiplied by the inefficiency of scaling a distributed system, especially on a network you don't control, may well be higher than the price premium of a mainframe.
OTOH, if "you" are like Google, and are paying commodity (or below) prices and have had operating your own hardware, as cheaply as possible, as part of your strategy all along, including investing in experts/specialists to keep it cheap, then mainframes wouldn't make much sense. Amusingly, it seems popular to emulate only part of the big players' distributed systems, the software, while ignoring the context that made it so useful for those players, the cheapness of the hardware.
[0] I'm using quotes because the terms are a bit imprecise in comparing mainframes, which, even traditionally, involved distributed processing. IIUC, the reason there's a "C" in CPU is from mainframes, which routinely relied on processing units that were not central. Other commenters have asserted that mainframes are horizontally scalable without detailing what that means.
[1] https://en.wikipedia.org/wiki/Amdahl%27s_law
[2] http://www.perfdynamics.com/Manifesto/USLscalability.html
[3] https://en.wikipedia.org/wiki/Fallacies_of_distributed_compu...