Stratus: Servers that won’t quit – The 24 year running computer
cpushack.com
cpushack.com
I can see how Google is better off spending money on software engineers than on hardware (thus, fault-tolerant distributed systems); but there's lots of systems that don't need Google-scale computer power, nor Google-style ultra-low-cost-per-operation hardware, but that really need to stay up - utilities, banks, every system at the center of a big enterprise.
Is the issue that modern software just isn't reliable enough to make hardware failures an important part of the downtime? Or is it just that systems like https://kb.vmware.com/selfservice/microsites/search.do?langu... are cheaper today? (I guess those actually work, right? The level of complexity involved is frightening...)
In the modern era we continue to have the same small number of civilization critical systems, we still have exactly one air traffic control system. However we have billions of, well, filler. Ringtones and social media and spam. Therefore rounded down 0% of modern computers are doing civilization critical work and safe secure systems superficially appear to be a thing of the past even if the number of systems and employees remain constant.
In 1960 the average ALU was calculating your paycheck and it was rather important to get that done on time and correctly. In 2010 the average ALU is a countdown timer in your microwave oven, its your digital thermostat, its my cars transmission processor, maybe its video game system or other human interface device (desktop, laptop, phone) but probably numerically unlikely. Things are sloppier now because they're sloppy-tolerant applications.
Flexcube now owned by Oracle (core banking, basically a general ledger if used to that extent) and VisionPlus (credit cards) are done in multiple languages, but the core is COBOL. If you're in banking you probably have a guess who I worked for.
These systems go deep into an organisation's nervous system, and in the case of Flexcube, for example, come from a prominent bank's nervous system.
Most in-hourse mainframe programming today however is on the middle layer providing APIs with JSON feeds, heavy lifting left to 3rd parties largely with workforces in India and sometimes China.
This is a world away from trading or investment management. While for a layman 'finance' is often a catch-all, banking IT and trading IT are different worlds. Java dominates investment/trading, with C# doing a bit but not much. C++ is mainly plugins for traders written for something not self-taught VBA that need to do things fast, but Java has largely filled that void supplemented by C#.
Edit: Around 5-10 years ago, a lot of IT work in banking was about unifying each country's code-base around a common core, the larger the bank the more tedious the task. Now it is all about AML and KYC, if you're interested in banking non front-office space. ML is a joke, quants do what quants do and win 50% of the time.
As zhte415 lets on, most of the people working on the mainframes were third-party contractors. I don't think the company even advertises COBOL and mainframe roles. When I was hired (as part of a graduate programme) the job description didn't say a word about mainframes. I had to raise a little hell to get into one of the mainframe teams, and it wasn't even particularly easy.
I mean, you'd think with all the stuff people say about how the old guard is retiring and those ancient systems will need people who know how to maintain them, they'd have been waiting for me with open arms. Not quite.
My hunch is that there's a supply/demand thing going on. It's not that all those big banks etc. don't need modern-educated software engineers that can tend to the ancient tech. They do. Except, those new generation soft. eng's don't really care about the ancient tech, and the corps can buy cheap foreign labour to tend to the mainframes. There's no supply and all demand is covered. So they advertise for the roles that they do need filled, which is to say, everything on mobile platforms, the web etc.
I'll narrow this down to one particular section of our consumer bank, because that serves as a fairly "clean" example, not much affected by outsourcing other complications.
We have a "core banking system" -- the computer that keeps track of the account balances for customers and moves around the bits that represent dollars. This is implemented on 1960s technology: ours doesn't happen to be in Cobol, but it is written in a different language of similar vintage. The system is pretty rock-solid: as an example, it has only a couple hundred milliseconds at most for all of the processing needed to approve or decline a credit-card transaction in real-time, and it has no difficulty checking balances, restrictions, and business rules to enforce that -- that isn't so impressive until you consider that even the slowest transactions need to meet that limit and we still have to allow for network latencies. We struggle to find qualified programmers for this system: the few who are actually skilled in it can command pretty decent salaries and benefits, and we'll search around the world to find them. The system is not abandonware: we'd love if it were written in something more modern, but a complete rewrite would be an incredible undertaking (this system developed over several decades, and that is difficult to recreate). On the other hand, we ARE looking at questions like how to get it to run in the cloud.
So that's all true, but if you look through our technical job listings, you'll mostly find us looking for Java developers, PostgreSQL DBAs, Angular web-developers, and other such positions. That is because the core banking system is a tiny portion of what we do. In my specific area, we have roughly 20 development sprint teams (and a few other support folks not on those teams). Of those, ONE team has core banking system developers on it (along with some developers with other skill sets). By that estimate, it makes up 2% to 5% of the work we do.
The fact is, keeping track of the balance in your account is only ONE TINY PART of what your bank does. We have to keep track of your personal information (email address, mailing address, login id, etc). We have to serve you up a website. We have to process scanned checks, send out marketing emails, analyze traffic to detect fraud, and hundreds of other things. The core banking system (and other similar systems in similar companies) may be written in 1960s languages because the existing systems are robust and well-tested, but for that exact reason, they don't require a huge amount of development work.
Just out of curiosity, what kind of numbers are you talking about?
> On the other hand, we ARE looking at questions like how to get it to run in the cloud.
Call me a masochist but that sounds really fun.
Unfortunately, that's exactly the kind of thing my employer does not want me to talk about.
> Call me a masochist but that sounds really fun.
Oh, it is. It's actually a really interesting project, and from what I've seen so far I think it may well be fully successful and go live with our actual customer records by mid-year 2017.
Oracle market a completely different code-base (which is fully Java), even single-currency code-bases, from those for larger banks, small third parties 'partners' providing implementation.
I do wonder what they use. I assume off-site computers aren't really an option for networking reasons, so Stratus, NonStop, z mainframes, etc probably. I wonder if they have a backup mainframe, or if the redundancy in one of those things is enough.
Nobody is allowed to rewrite the nuclear power plan software. My first job was at Westinghouse Nuclear Division. Those Fortran libraries were off-limits in terms of changes.
WTF? What is energy balance circle management?
In short, yes. They always have two machines running ready to swap in. I have heard some also run multiple live mainframes to check that the results agree, and discard and restart calculations if they don't.
If you have your service spread across 10K servers, well, it doesn't matter if some go down.
Instead of making them all 100% super fault-tolerant for a crazy expensive unit price, you can make them relatively cheap and replace them when they fail.
Most people don't value 99.999 vs 99.9999 reliability as worth millions of dollars per system per year. Space shuttles may disagree, but not billing platforms.
For something like the elevons, control was accomplished by connecting 3 actuators to 3 of the channels, with voting being accomplished by physical force - if the 3 computers disagreed, the 2 would overpower the one. (Things like thrusters used electronic voting, close to the thruster itself.)
This seems closer to modern architectures with multiple computers, than the mainframe idea of redundant hardware.
"At 12:12 GMT 13 May 2008, a space shuttle was loading its hypergolic fuel for mission STS-124 when a 3-1 disagreement occurred among its General Purpose Computers (GPC 4 disagreed with the other GPCs). Three seconds later, the split became 2-1-1 (GPC 2 now disagreed with GPC 4 and the other two GPCs). This required that the launch countdown be stopped.
During the subsequent troubleshooting, the remaining two GPCs disagreed (1-1-1-1 split). This was a complete system disagreement. However, none of the GPCs were faulty. The fault was in the FA 2 Multiplexer Demultiplexer. This fault was a crack in a diode. This crack was perpendicular to the normal current flow and completely through the current path. As a crack opened up, it changed the diode into another type of component ... a capacitor.
Because some of the bits in the signal are smaller than they should have been, some of the GPC receivers could not see these bits. The ability to see these bits depends on the sensitivity of the receiver, which is a function of manufacturing variances, temperature, and its power supply voltage.
From the symptoms, it is apparent that the receiver in GPC 4 was the least sensitive and saw the errors before the other three GPC. This caused GPC 4 to disagree with the other three. Then, as the crack in the diode widened, the bits became shorter to the point where GPC 2 could no longer see these bits; which caused it to disagree with the other GPC. At this point, the set of messages that was received correctly by GPC 4 was different from the set of messages that was correctly received by GPC 2 which was different again from the set of messages that was correctly received by GPC 1 and GPC 3. This process continued until GPC 1 and GPC 3 also disagreed with all the other GPC."
[1] Adapted from https://c3.nasa.gov/dashlink/projects/79/wiki/test_stories_s...
There's a bit of debate as to what would have happened if it ever took over.
Years ago, I had a customer that had a Tandem which was quite exciting to me as I had yet to encounter one up to that point of my career.
Imagine my letdown when I eventually discovered that the Tandem was used for FTP. ¯\_(ツ)_/¯
For highly critical systems that is an issue but it is being adressed. Take as example the space shuttle or the Airbus (3/2) primary/secondary flight computer [1] where the backup is being build by different companies with different processors ...
[1] https://ifs.host.cs.st-andrews.ac.uk/Resources/CaseStudies/A...
Same brand, same batch, operating time, running temperature, even same vibration does trick. :(
And if it's a software problem, then again the redundancy isn't helping you.
Basically you haven't succeeded in getting 5 independent tries in the problem -- the failures are highly correlated.
It's the same reasoning that led to the financial crisis -- you have a CDO or an MBS that has a bunch of different obligations so you think you've diversified away risk, but in actuality there is one big factor that causes everything to fail at once.
Rather, the issue is that individual machines are cheap enough that you can have machine-level redundancy at a much lower cost than trying to engineer a single machine with redundant components.
The issue with machine-level redundancy is that they tend to run their own isolated operating systems (which is great from a fault-tolerance perspective, but is rather wasteful from a computational perspective). Operating systems like Plan 9 were billed to help bridge this gap and make clusters of machines communicating over 9P feel like a single unified whole, but they never seemed to catch on (besides maybe the concept of a Beowulf cluster).
Sometimes people developed redundant but not non-stop systems. I talked with an VP of Citibank when working at NetApp and they had a number of systems which ran on schedules of alternate days, so one would process transaction records for a while, then another would take over and repeat. They had three identical systems where one was eseentially a hot standby for the other two, and new versions of code would be deployed on one which would run the same transaction records and they would check for the same output, so they could do a 'walking' upgrade of software. Back when Tandem's and big Sun iron ruled the roost those machines were too expensive to have an extra one which was essentially a spare. These days however its much more economical to do that.
Source: I work with several ex-Stratus employees who have told me some of their war stories.
The stratus server I worked with in around 2002 was, I believe, one of the first to move towards standard x86 configurations and Windows OS.
They did so because, no matter how reliable your server, executives were jumping on Windows as an application server bandwagon and developers were targeting it.
There were several driver and firmware bugs that would BSOD the whole server, and bring the entire fault tolerant OS down.
Moreover, storage was a hack job. They couldn't make hardware RAID cards work in that configuration, and Windows software RAID was a joke, so you were stuck with Veritas storage software. Twice we applied Windows Service Packs only to find it would BSOD due to a Veritas software incompatibility. If it was a more average server I'd apply said update to a test server first, but no one could afford a test Stratus so you were always testing in production, which doesn't help reliability.
Here's a picture of it. It became my homeserver for a few years but frankly it was less reliable than a desktop.
You could more easily justify the cost when the upcharge wasn't so huge.
And, of course, we all got better at reliable distributed systems with commodity servers. So aside from the cost gap, the uptime gap was closing as well.
In the mid-90s, data centers full of commodity hardware would have been cost prohibitive.
Your comment about software, though, is a good point. The last relatively common OS that has that kind of uptime is OpenVMS.
It's possible to get close enough with off the shelf hardware and good system design, so nobody wants to pay 5-10x more for special hardware.
I think the issue is more about networks never being reliable enough for such reliable computers to ever make much sense. A data center simply cannot give you five nines availability over the internet, so it doesn't matter how reliable everything inside is, you would still need geo redundancy and all that fault-tolerance in a distributed system.
I almost couldn't bring myself to shut it down to do the upgrade. I mean, I saw servers with 100 to 300 days uptime routinely, but not one that had never been interrupted for over 1000 days....
One day, the shared mount on a ton of systems around the company just stopped working. The master node (these have internal redundancy at the network, controller, shelf, and disk level) had failed over to the redundant node, which after time had so many disks fail without being replace that it went down hard. They had to work with NetApp support (at great expense) to get it back up, which took a few days. The system had been attempting to warn the admins as best it could of the impending doom for the better part of a year, and no one noticed.
The moral of this story is that you can have big-dollar redundant systems, but if you don't also have some dollars going to the care and feeding of them, it's wasted. The other lesson here is that such systems don't exist in isolation and are only as good as their weakest link, which in this case was a now-unreachable SMTP server. Lastly, although the LCD displays were trying their best to also signal this fault condition, people who would walk by these physical machines were used to this, so they didn't bother to actually read the errors. Do that at the peril of anything dependent on that system.
This is the area were inverted Unix philosophy applies: Give me good news as well. A weekly mail, that backups (or the filer or RAID or whatever) works A-OK. If the mail stops coming, something broke.
Agreed. Another moral: Don't believe anything works until you see it with your own eyes.
It would be interesting where I am since you skip local crime and go straight to federal offense. Kept me away from booze in high school since its not minor in possessions, its federal bootlegging. I can only imagine what trying to enter a server room nets a person.
By the time it was being looked after by the third party maintainer they'd just leave us a few disks and a couple of fibre channel controllers onsite and we'd get the DC remote hands to swap them out. The faulty parts were then boxed up and couriered back to the maintenance company all on their dime.
There was not an awful lot of involvement on our end except for authorising access to the DC when an engineer needed to attend.
None of this was massively expensive (and this was 24x7x365 cover) in the great scheme of things given their mission critical status.
You have essentially described IBM's business model. You pay more, but if something goes wrong someone shows up.
Yes. They are expensive and deviously good at it. Or, at least they used to be many years ago.
Many years back I worked in a tiny town (1000 people?) four hours from the closest depot and we even got same day service. This cost a king's ransom. But the AS/400 ran most of the business...
I think the bigger IBM z machines do allow zero-time upgrades, but the iSeries (new name for AS/400) is basically a reboot the machine affair. I've never heard of Sysplex in regards to the iSeries.
They sent me two DVDs with 12 pages of instructions. The unnumbered first page is release notes and a notice that physical media deliveries would be an average of 5 to 9 business days on Jan 31, 2007 and beyond. (odd notice for a 2016 document)
Page 1 has my customer information with all sorts of tracking numbers. It also include the number of selected files (1,360), Kilo-bytes (sp) of data (3,963,529), fixes (3,394 fixes with 815 Kilo-bytes of data were superceded), etc.
Pre-install instructions finally show up on page 3 and the big warning about f'n it up shows on page 4. Page 6 brings us to the actual install which is written to be typed exactly with exact detailed results of each step.
to give an example
2. Load the image catalog into the virtual optical
device using the following command:
LODIMGCLG IMGCLG(ptfcatalog) DEV(OPTVRTxx) OPTION(*LOAD)
3. Type GO PTF and press the Enter key
4. Take menu option 8 and press Enter key
Took about an two hours counting the the 45 minute reboot. We have a very old and slow machine (need to buy a new one in the next two years, we are at a decade of use now) so reboots are a bit slow.https://www.vmssoftware.com/about_faq.html#faq1
They didn't buy it. HP is too greedy about the lucrative revenue & maybe patents they get off VMS that made me wonder why they killed it in the first place. Instead, it appears (not certain) that VMS Software Inc gets some OEM-like license that puts them in control of it while HP still owns it, can license it themselves (probably existing customers), and likely gets a portion of VMS Software Inc's licensing revenue. So, they've pushed off almost all responsibility for it onto another company without actually selling it or losing all the revenue.
Still a good development for VMS customers. Both legacy and people who will suffer through archaic stuff to reap benefits of bulletproof clusters. I'd be happy to. :) Only problem is it's basically got no security attention over the years with plenty of zero days or configuration issues lurking in there. If I deployed in a company, I plan to put network guards in front of it to absorb attacks while converting the traffic into something easy to process. Ideally, a PCI-based guard that forces it to only talk to the application's memory instead of rest of system. Plus safe-coded apps (Rust, Ada) and API wrappers to try to spot BS. That VMS supports cross-language development will probably help.
EDIT: The weird thing is they finished a port to Alpha ISA before they got to x86 ISA for migrations. I thought Alpha's already migrated to Itanium. Didn't see that coming lol.
http://h41379.www4.hpe.com/openvms/whitepapers/high_avail.ht...
Stratus has been doing that since the 90s, though :-)
It was very common to have many months of up time, despite a lot going on (virtual machines, games, etc.).
Continuums were working till 2003 when they were replaced with x86-based Stratus servers, which were double or triple redundant, depending on the model. They are still running at the RTS Exhange which is now part of Moscow Exchange (formely MICEX).
The only problem with x86 models, at least the models produced back in mid-00 was that the backplane with synchroniztion clock was not redundant and if the clock failed the whole server went down.
Also, how does this compare to the oldest running satellites?
[1] http://www.computerworld.com/article/3162416/data-center/boo...
Its very cool to see such reliable systems in practice.
They also offer a lower end product based on two standard servers where all the redundancy is software based with virtual machines.
It's clever kit and looks after a lot of critical infrastructure where downtime isn't an option.
And this is there software based solution which runs on commodity hardware: http://www.stratus.com/solutions/software/everrun-enterprise...
I only work with the latter but have heard the hardware offering is bulletproof.
This isn't true at all. The cloud offerings are more like traditional, high availability such as fail-over clusters. They might compete with VMS clusters by now with the right components. Maybe even exceed them except for longevity. Systems like NonStop and Stratus are fault-tolerant systems supposed to get five 9's availability with [hopefully] imperceptible moments of failure that [hopefully] recover immediately. I wouldn't trust a cloud platform for that. The latency alone would probably prevent the solution from matching a local, NonStop cluster.
But impressive kit, not cheap but at the time was one of the best choices outside the AS/400/IBM route and indeed Lloyds did purchase a few of these systems.
More so given they lasted this long in use and even today I've had best in breed systems have faults that were not iditeified as soon as possible, recall some fancy IBMkit having a dieing PSU and could smell the electrical burn faintly without any diagnosstics flagging up any problem for that server to fail in a few days time.
But today we tend not to have one large all singing and dancing stable system and often cheaper to have redundancy thru load balancing and with that able to pull a server out of a work load cluster. This along with Assured messaging systems (MQ etc) you negate so many hardware faults that can and did in the past put a halt to work.
A somewhat more realistic way to build it outta real hardware involves each processor runs every 1 outta X clock cycles. Reading between the lines of the story WRT "down clocked processors" I suspect this is what its doing. This also detects weird power spike problems where lightning miles away means every processor running is going to run 0xFFFF whatever opcode that happens to be, but a couple ns later the spike is gone and the other processors are back to normal. Or RFI/EMC, etc. This is all very nice other than a significant hit to performance if you run a large number of processors. Then again if your primary figure of merit for your system is extreme reliability and not raw MIPS...
It would be interesting to hear if they run dual port RAM to get around the jitter XOR problem and have some alternative way to detect synchronous behavior. Dual port ram is weird but COTS. Triple port ram is not as COTS.
Past that I dont recall much.
how does this work? how does the system know which pair is the faulty pair?