Moved a server from one building to another with zero downtime
reddit.com
reddit.com
This anecdote is an amazingly good story for telling at the pub over a few beers. It's a terrible story for a strategy.
If this is a mountain, my molehill is that one night in the late 90s, I got paged cause the SMTP outbound server was overheating. At midnight I drive across sleepy NH backroads, and stopped at a Wendys to get a chicken sandwich and iced tea, for the caffeine.
When I got to the server room, I pulled the 2U Dell server out of the rack and discovered the CPU cooling fan had seized up. Mind you, this is a New Hampshire data center in 1999, and it has a filing cabinet with manilla folders, and carpeted floors. This thing was never prepared for any disasters.
A half hour later, the SMTP server was up and running cool again.
I greased the fan with the mayonnaise from my sandwich.
I gave up on book 7, so yeah, I tried sticking with it.
Unfortunately there is little plot advancement, perilous situations more contrived, and needless exposition/filler the norm.
As in, if I’m cleaning and have nothing else interesting to listen too.
Some occasional good laughs. Dinosaur holding a plunger badge. :)
I believe the "Nope" sentence was something akin to "{character} thought that {thing} because {thing}."
Jesus Christ. Would it kill you a little to show instead of tell?
(And lest people believe I'm not a fan of some good * opera, I'm not ashamed to admit I've read my fair share of BattleTech, Barsoom, and even Lost Fleet, among less highbrow works)
It sticks in my memory because I remember being somewhat annoyed at the writing level hitherto, reading through that sentence, realizing what I'd just read a few sentences later, going back to double-check, then chucking the book.
And don't get me wrong, I've got space for some crappy writing in my sci-fi (looking at you, Foundation).
If I microwaved it from cold, it would break almost instantly.
Here's the ingredient list for Wendy's mayo: Soybean Oil, Water, Egg Yolks, Corn Syrup, Distilled Vinegar, Salt, Mustard Seed, Calcium Disodium EDTA (To Protect Flavor)
Over time, the server would expand to be 4U rather than 2U
That’s how you get ants.
Except that it had been 5 years since the last maintenance in this place and it was a protection panel for a large synchronous generator in a power plant.
After you make a heroic temporary fix, please, ensure the permanent fix is applied later!
I've known people who would, depending on exactly how bad the failure was, outright refuse to apply temporary fixes precisely because they didn't believe that the business would fix things properly if the issue wasn't forced. And having watched how that particular company handled things, I can't say that they were wrong.
I also think the converse holds (All permanent fixes are temporary).
https://www.google.com/imgres?imgurl=http://i.imgur.com/TtFo...
...it looks like Google is referring to imgur to reddit? In any case, here is the link that it goes to.
https://old.reddit.com/r/funny/comments/26vx0x/handy_fuse_re...
I'm more worried that the oil won't get to a high enough temperature and thus won't polymerize, so it'll flow out and ruin some other component, or go rancid, or something. Thermal paste won't move on you. Oil will.
https://www.pugetsystems.com/submerged.php as an example.
Pretty much any Solid-liquid-solid or solid-solid-solid interface will be better than solid-roomTemp&Pressure gas-solid.
The whole point is to conduct heat better than air, and most things will.
If the story is true, the client is a stereotypical know-it-all small business owner who gets by on bullying. You see them frequently in businesses that pay low-skill workers a small premium that is hard to replace. (ex: cleaning services, pool guys, mechanical contractors that do low-end maintenance work, etc)
As a contracted SME, taking a job like this is dumb. The chances of failure, where "failure" == the server going down is high, and the customer will just stiff you.
It’s so frustrating and stressful.
Walk the length with the appliance and verify there's no dead spots, then just hook up to a power supply and get things done.
When you consider you're getting like 1000 of the smartest tech people in the world to manage your infrastructure for you, for $5/month per user, it's really such a no-brainer. If people are too stubborn to see that, or want to waste time trying to do it better themselves "because it's cheaper", with redundant power, OS patching, zero-downtime changes/deploys, proper capacity planning, proper redundant connectivity, provisioning the right network around it, physically securing the server room, ensuring things are properly cooled, not wet...I could go on forever, and this isn't even the main focus of the business...I'm sorry, but they deserve go to out of business.
I had a very stubborn client once who ran a hotel chain. Won't say what or where, but I wasn't surprised when their random "security through obscurity" VNC server got compromised. I wasn't finishing migrating to the new PCI-DSS compliant system we built, either, so there go 5000 credit cards "encrypted" with some sweet rot13-level bullshit in Turbo Pascal I cracked in about 30 minutes with no code access.
Agreed though, this particular customer disqualified himself as soon as he said he wont pay if the server goes down. He should have offered a big bonus if the move succeeds without downtime.
See my comment elsewhere on "the customer is always right."
Sure. Just, there are some verticals where charging a "positive-ROI" amount gets you no business at all, because all the potential clients in that vertical are businesses that operate on such razor-thin margins that they don't actually have the cash-flow to pay for the extreme requirements they also have. They've been getting along until now purely by begging/tricking/manipulating people into doing negative-ROI one-off tasks for them. If forced to get contract all the services they need out on the free market, their business would cease to exist.
(Therefore, you say, they should cease to exist. I'm not arguing!)
if you do, you're just selling dollar bill for 80c. You may drink growth-kool-aid, or someday-monopoly-hope or VC-subsidizes-business.
In the end, somebody pays for it either from stupidity or hope.
I.e., a job paying a lot of money is "return" on investment; but so is a job that pays a lot of experience, or satisfaction. Meanwhile, a job that takes a lot of time to do requires a larger "investment" (in terms of the set of BATNA jobs you could otherwise be doing with your skills); but so do jobs that don't take very long, but are more physically demanding/emotionally draining, or where the client keeps jerking you around by changing the requirements. If working the job makes your utility go up (i.e. makes your life better), it's a positive-ROI job. If it makes your life worse, it's a negative-ROI job.
So, let me restate: in some verticals, while the clients do pay a livable time-and-materials wage for the region and experience level... they expect far more labor (and far more highly-skilled labor) for their dollar than that dollar "should" be able to get them, such that the net utility from anyone taking that job is negative, even if they're "making money."
These businesses get by mostly on exploiting young IT ingenues who don't yet realize their worth, and so are willing to give a try to e.g. "program a social-networking app" for $2000; or to e.g. "set up a PXE Linux automatic memory-testing boot env and Windows system-image deployment pipeline" for $300.
(These are both things I naively got myself into when I was younger. Amusingly, I was in over my head on the social-networking app, but managed to complete the hardware qualification+deployment env just fine. And my only major obstacle for the social-networking app, was that the client insisted that it needed to be hosted on some free ISP hosting they had, and thus needed to be written in PHP4. Which I didn't even complain about... it was just a language I wasn't very familiar with, so I tried to learn it while working. Sigh.)
In other words, these companies try to get newbies to do senior-level work, paying only newbie costs. Sometimes, surprisingly, that works! And they survive entirely because it sometimes works.
It's basically the "my nephew knows computers, he can certainly repair mine" mindset applied to corporate planning.
So, 1 hour billed for 5 minutes of downtime, or 40 if you want absolutely none. Happy to do either, but highly recommend the former. 99% of people pick the cheaper recommended option.
In this case, I would've tried to put the server on WiFi which would seem like less a hassle for me. Equipment acquisition cost billed to the customer.
If you offer a client an option with a 1% chance of a very bad outcome and they understand and accept the risk, sure.
But if they have not understood and accepted the risk? Implementing something you already know the client misunderstands ain't always a smart move.
At the end of the day, it doesn't really matter if they listen/understand or not, as long as you documented telling them or it makes it into some legal document and they still pay the bill.
I don't know, that sounds too close to encouraging the attitude of "it's not worth it, I'll just take my normal pay". 10x vs. 0x is a significantly stronger incentive than 10x vs. 1x.
While I agree that it's difficult to predict exact credit risk based upon customer personality, explicit threats not to pay you like exist in the story -are- a bit of a signal that there may be a risk of nonpayment.
Since the customer is never going to know the awesomeness of it, it's really just for yourself.
It was ok, just the power cord and a fee RJ45 torn. No serious damage besides downtime.
Build yourself a home lab and learn how systems work. Figure out what is really running your code. Learn how to resource optimize.
Don't end up only being able to work on webapps and small datasets that fit comfortably in the cloud.
5 years ago I left a company running loads of enterprise software and web applications for other companies in a couple of data centers. We had well over 10.000 physical machines. When I left, about 90% of the workloads were virtualized. Running on bare metal is the exception today.
Sure, its cool to have hot-swap HDDs, hot-swap PSUs, redundant network cards and a remote management board but none of that is rocket science. You can learn it in a week if needed.
Networking is a much deeper topic yet almost nobody would recommend you to set up a couple of switches and hardware firewalls in your home network. Today you expect that the hardware just works. In the real world it is extremely complicated to correctly design and run networks, but you aren't going to experience any of these problems in a home lab.
At the end of the day, you need the guy who makes it "just work". That is a valuable skill as well.
Also knowing how your application works when you just yank the power cord on half your VMs teaches you a ton.
-"Why? I don´t know why. It just works."
-"Is Hellmann's ok?"
IT person documents that Hellmann's is preferred.
Essentially it's a medium sized double conversion ups, with a really high quality sine wave inverter, and some electronics that can match phase with a live 120vac 60Hz circuit. And a tool kit which consists of the insulated electrical hand tools needed to do a midspan removal of the cable jacket and splice into the wires in an ordinary PC power cable. The person using it is of course supposed to be trained in advance, and competent at the process of attaching the UPS to the live circuit.
I think the thing I saw is also meant to deal equally well with a commodity x86 PC built from parts, or an Intel NUC size thing, or a corporate desktop machine with proprietary internal wiring like a slimline Dell, Lenovo, HP, etc.
Dealing with AC power isn't really that dangerous if you're careful.
So long as you’re not earthed https://imgur.com/gallery/B2c5FfD
Lesson learned: electrical wiring is like a gun. Always treat it like it's on, and if you have to do something would be unsafe if the wiring is energized, make damn sure it's de-energized before proceeding. When you're working in that mindset already anyway, flipping the breaker for something as simple as swapping a switch/outlet hardly has any benefit.
1) Treat every wire as if it was hot. Even if you know it's not. 2) A good electrical connection must first have a good physical connection.
Not sure why that second rule sticks with me :) but there has been more than one occasion when I'm fairly sure the first rule has saved me from a bad shock. And you're right - treating the wires as if hot means you can actually work with hot wires for a lot of simple things.
I still turn off the breaker though :)
The wire nut is only there to stop the wires loosening over time and provide some basic insulation. It is not there to actually attach the wires. When you twist your wires together, they should be attached well enough on their own that you'd be comfortable throwing a piece of electrical tape over them to stop them shorting to the box and leaving it as-is (but don't do that). If the only thing keeping them together is the wire nut and you being very gentle when you manipulate them back into the box, they're not actually connected.
The poor physical connection creates a poor electrical connection. A poor electrical connection has resistance which creates heat. Heat creates fires. Even better after a few years when enough traffic has driven past your house and enough people have moved around inside of it and the wires have wiggled to just barely in contact so occasionally when someone walks down the hallway the lights will all flicker as the wires create some pretty electrical arc light shows, adding carbon buildup to the wires and further increasing the resistance and heat concentrated in the one tiny point of the copper where they're still sometimes connected.
No reason at all for this rant. Definitely not a real example at all. Definitely didn't waste an afternoon with a toner, a drill with a pilot bit, and a borescope to hunt down the six octagon boxes someone had sealed into the basement ceiling hiding away some of the shoddiest wiring I'd ever seen. Nope.
Ground is ... all the metal bits in the bathroom and there's earth leakage happening somewhere.
The safe work procedure is then: get the shower to the desired pressure and temperature before you get in / while you're still wearing your shoes then try not to touch the taps while you're in there.
But don't tell the guests cos hearing them yell "FUUUUCK!" is amusing.
Bonus points if they pass out from the shock and knock their head on the way down.
Caring bunch us Aussies.
I'm not a nut that does everything with the power on--I kill any branch I'm working on and double and triple check with a non-contact voltage detector before I stick my fingers into anything (which saved my bacon the one time when the hot from a different branch of the same phase ended up connected to a neutral wire for a plug with no connected ground leaving it showing 0V on a multimeter in any configuration and still being live with the breaker off; that house was a mess). However our current dwelling has no main cut-off for the power. If we wanted to turn off power to the panel we'd need to get the power company out to pull the meter from the socket.
In a mostly full panel the bus bars are pretty much completely covered by the breakers anyway. You'd have to work pretty hard to come in contact with them. And the wires you're working with (besides the ground) are insulated anyway so no issue if they brush up against something.
The only thing that's _slightly_ butthole puckering is chasing the uninsulated ground wire through the panel down to the neutral bus.
And yeah, done without gloves because weighing "safety when I make a mistake" versus "greater dexterity so I'm much less likely to make a mistake" I prefer the latter. The protection is rubber soled shoes and keeping one hand tied behind my back so the electricity has no path through me.
He did this without notifying the power company, so those supply lines were hot with 240V residential service. The weather shifted and a light mist started falling before he was done. Like another poster above, I was thinking I need to be ready to call 911, but wanting to be far enough away not to be hit by splattering metal or any surprise voltage gradients in the soil.
But, careful work habits and some tools that happened to be insulated anyway, meant that I was never bridging two different potentials. The job went flawlessly and I only noticed when I plugged the outlet tester into it at the end, expecting to go turn the breaker on and come back and look at the lights... but the lights were already lit up.
For one project we had large banks of ultracapacitor in a cabinet. Fully charged it was around 1200 VDC. This thing was in the prototyping stage, and we were testing a control system on a Saturday morning.
So we charge it using a large AC/DC converter, fully charged, everything worked beautifully. We start a discharge cycle converting the DC back to AC. Uh oh, it starts pulling way too much current. Flames start to shoot out of the AC/DC converter. Fuck. BANG. Fuse blown.
We assess the damage... the AC/DC unit is totally shot. And someone (me) is going to have to analyze what caused the failure. Otherwise everything with the capacitor cabinet seems okay, but the thing is still charged to 1090 VDC and the fuse is blown. Check with the mechanical engineer that designed the cabinet. Turns out the fuse can't be changed (can't be accessed) while the cabinet is charged and the cabinet can't be discharged because the fuse is blown. Well that isn't good.
The only thing we could do was discharge it into a load bank (think large toaster) by connecting something directly to the copper busbar live at 1090 VDC. So one of the commissioning guys volunteered. He put on some high voltage gloves, stood on a plastic mat, and connected some jumper cables someone had in their car to the bus bar. He stepped back and someone else threw the switch on the load bank and it discharged without incident.
There were some design revisions after that.
However, there are anti-jigglers too that lock the machine when any new human input device is plugged in.
You could have a list of known USB device IDs you trust, and if a newly plugged in USB device wasn't on that list you could lock or power down.
They didn't go so far as to cause alarms on unknown device ids, but devices would just not be mounted if they were not whitelisted.
This was during the windows XP era when it seemed there were an endless number of security problems related to usb devices, no matter how good the group policy and registry settings pushed via active directory membership were.
The safeguard doesn't need to be perfect, it just has to be good enough.
While these second order effects are immeasurable, they are quite tangible in my personal experience.
Attempting to authenticate USB devices is a very hard problem — a sufficiently advanced attacker can spoof manufacturer and device IDs, even if you lock things down to prevent anything other than a keyboard or mouse it's possible to send keystrokes to open the wrong website, there's always a chance of an exploitable flaw in your USB stack, etc. — but anyone diligent can be paid to walk around every week checking to make sure that a seal is solid and the tamper-evident stickers have the same serial number as listed on the inventory. There is a real value in having things where the failure modes are obvious and intuitive.
Here's a current story:
Someone ordered the wrong desk phones at your large company?
1.) Assemble your crew. Go to various departments and recruit non-technical people.
2.) Task them with disassembling 1000 desk phones.
3.) Hot glue USB port on phone shut.
4.) Reassemble 1000 desk phones.
(Due to various disability acts they can't really do it either, as the employer must provide their staff with hardware they require, e.g. ergonomic keyboards and mice)
(For your parenthetical I should clarify - it wasn't the case that it was impossible to whitelist other devices, it just had to be done on a case-by-case basis. I.e. you would call IT and say "Jen from accounting at machine foo123 needs her new ergonomic mouse to be recognized" and they would remote in, tell Jen to unplug and replug the device and whitelist that exact USB device id on that exact machine.)
that...or the USB port is permanently blocked (saw that when I was at a finserv years back: all USB ports (except the one the mouse plugged into) were epoxied
HotPlug must only work in countries with terribly designed plug outlets like the US and Canada. Our NEMA 5-15 plugs are live when the plug's hot (electrons be here) and neutral (return to sender) blades are still visible. I don't think this device could work in the UK I'm not from there but I think their plugs can't be live with exposed plug blades.
https://www.cru-inc.com/products/wiebetech/hotplug_field_kit...
Unplugging just enough to expose the prongs is risky because the point where contact is lost will vary from receptacle to receptacle.
Chances are things are plugged into a multi-plug hub anyway. European homes are especially lacking in sockets in my experience.
I could pretty easily write a script that forces my machine to reboot and do all manner of other things if some sort of network change is detected.
There was also one that would use the motion detector to try to detect if the device was falling, and park the hard drive heads before impact.
Then there are splitter clamps.
The most common ones are used (low voltage) on cars and motorcycles, they look like these:
https://www.mesconnettori.it/index.cfm/it/ricerca/?flbr=ruba...
But there are professional ones, suitable to 110 or 240 V example:
"With the CRU WiebeTech HotPlug you can transport a computer without shutting it down.
"The HotPlug allows hot seizure and removal of computers from the field to anywhere else. The HotPlug's patented technology keeps power flowing to the computer while transferring the computer's power input from one A/C source (such as a wall outlet or power strip) to another (a portable UPS) and back again.
"We created this product for our Government/Forensic customers, but it has IT uses as well. Need to move a server without powering it down? The HotPlug can do it.
"It's great for digital forensic investigators and techs who can't risk losing access to data on a running computer. With many computers now employing full-disk encryption, shutting them down poses the risk of having to crack a password after moving the computer to a lab for analysis, which can greatly increase the time and expense of an investigation. When combined with a WiebeTech Mouse Jiggler, you also won't have to worry about the computer entering password-protected screensaver or sleep modes."
If I were to do this, I would try and find a secondhand server that already has similar protection built in, so if anyone asks I could say "I did not even know it came with this feature".
Having a good security architecture is not obstruction of justice. Doubly so if the data is still accessible to you after the failsafe is tripped. All you've done is prevent their ability to access the data before informing you of the existence of the warrant, using access mechanisms that - to you - are indistinguishable from an unauthorized access attempts.
> "I did not even know it came with this feature".
A documented threat model and security policy that justifies physical tamper protection and pulls inspiration from FIPS is a much smarter legal strategy than perjury. Consult a lawyer.
A decade later I look back with deep surprise that we didn't think to abstract out the service instead of the hardware. I don't know how many of those behemoths are still being bought, now I work almost exclusively with small server instances that can come and go on the fly. Micro services and AWS have taken five-9s in a different direction. I frequently think of Sun as a failed Hephaestus, in a Christopher Nolan film he would be brilliant but could only turn out clumsy tools because of his deformity, he hates the things he makes so he throws them away before completion. Men find these cast-offs and temper and refine them.
He worked a night shift and i use to go hang out with him in the noc and download movies (residential bandwidth was not what it is today). Odd he nor i ever get in any trouble for that heh.
1. Shared state storage systems that supported replication were rare (I think Oracle and Informix maybe?)
2. Virtualization software was in its infancy (did SunOS have something before Solaris?)
3. RAM and hardware were waaaaay more expensive, meaning you often had to buy more pure metal just to answer questions fast enough
At least that's my take on it based on my dim faded memories
In the mid-2000s, enterprises were (and in many cases, still are) running proprietary software with proprietary RPC protocols that had no available source code or other means of modification, and most had no support for application-level high availability, access control, or any other operational quality-of-life feature that people take for granted today. Rather, that functionality was handled at the infrastructure level, through things like the aforementioned Big Iron.
The world looks different today, but those machines made sense for the environment at the time.
AS/400's were capable of that in the 90's (possibly the 80's as well). Heck, they'd call IBM for replacement parts on their own. You'd show up for work and there'd be an IBM guy waiting to be let in. He'd swap out a part with no downtime, and be gone. I've seen machines with uptimes of over a decade with zero on-site IT.
This sounds like something they'd make up on NCIS.
not a novelty. the true bigiron vendors (not sun) had been doing this for decades.
mainframe reliability puts the upstart unix systems to shame.
this is actually so ancient it’s hard to find docs. here’s something from 1976. (the report is 1990 but the hardwares dates to ‘76). https://www.hpl.hp.com/techreports/tandem/TR-90.5.pdf
What if a technical failure happen? What if there's a fire in the server room? What if there is an earthquake and the building collapses? What if... many things can happen that can result in a long, long downtime with this tactics.
If uptime is so crucial, the system should be setup in such way that moving one server should be a peace of cake, not a spec-ops mission.
If your mainframe had only one CPU, you did have to turn it off in order to service it. But you could upgrade the OS without turning it off. While they aren't cool tech now, mainframes are a marvel of hardware engineering.
(archive.org link because ibm.com apparently isn't hosted on a mainframe.)
Some circuits might average 5 to 7 nines of uptime over a year, but the next year is dump truck time... You can never truly be certain.
In the past power supplies and spinning disc hard drives would fail much more often.
It’s basically a solved problem, outside of extremely mission critical, 5 nines kind of stuff, that we all forgot because of AWS.
HN ran, and may still run, on a single bare metal server.
Also, that doesn't cover other problems mentioned here, like natural disasters, ISP problems, etc.
Also cost does come in play as well. Multiple physical links in would be very expensive for what sounds like internal services. Likewise a natural disaster might cause bigger issues to the company than those internal services going down. They might still have offsite back ups (I'd hope they would!) so at least they can recover the services but the cost of having a live redundancy system off site might not justify those risk factors.
The customers requires are definitely unreasonable though. I'd hope those systems are regularly patched, in which case when is downtime for that scheduled and why is that acceptable but not when you're physically moving the server? I doesn't really make much sense; but then "not making much sense" also quite a common problem when providing IT services for others.
In general, we don't know much about this case. It's a post on Reddit, might not even be true. As is, it doesn't make much sense, but we don't know all the details, so maybe we jumped to conclusions.
Mainframe is not just a server. You can hot plug RAM on these things.
Though it was likely easier to find than that Novell Netware server that was sealed behind some drywall, with only a stray network cable leaving any clue as to where it was.
I once only half jokingly suggested finding a missing data closet in a two million square foot distribution center by pinging a known IP from three or four aggregator switches across the building and triangulating the location on a floor plan. Sadly the people crawling around the ceiling found it before I could put my idea into practice.
Good job they found it!
Turns out a door was closed and a new one built to a hallway to another hallway and not properly labeled on the updated drawings. Had one of the boxes running a conveyor belt not have died, we'd never have looked.
If a function is super critical to business, it also deserves to have some thought put into the blast radius of its failure.
The sort of places that would insist on rolling a live server 700 ft across a parking lot probably don't have any real disaster recovery plan.
I've came across old AWS account (startup have been using AWS for the longest). All the network traffic or VPN goes through a single instance with 3 years of uptime.
I bet HN wouldn't do a 10 hours high-risk operation for moving their servers because they can't afford an outage. (But well, running stuff on a single bare-metal server is expensive enough that even if they could, I expect they don't.)
What would that company do if a pipe broke inside the datacenter? Besides, if you never restart your servers, you are guaranteeing that the one time when the power goes off on the entire city, they won't come back online.
HN is probably not business-critical and could probably affort a 10 hour downtime without much hassle.
While I don’t think they informed them of this in good-faith, it is a nice heads-up. In this case, it meant Consultant2 consulting RefusingConsultant that probably knew the IT better.
I hope there wouldn't be a correlation, but I wouldn't be all that surprised if a somewhat loose one was found.
Single server and "can't tolerate any downtime" are mutually exclusive.
There's SMART for disks... what else?
HN also has downtime fairly often.
You can see the common sense ship has sailed.
Whether it's smart and good for your business/reputation is a different question.
And then charge for those 5 hours for good measure.
In general, this stupid trend of wanting 0 downtime makes no sense to me. If you're not NASA, police or other emergency service you 100% can afford a few hours of downtime with scheduling it be forehead.
Second thing is having multiple machines for server. In theory it might help in increasing the availability but in practice I haven't seen any random issue due to machine which occurs just based on probability. I think almost all failure modes that exist, they are correlated between machines. eg suppose you have data loss on one machine, you could more likely than not, blame it on code and it would be similar across machines.
Re: multiple servers. Power supplies fail, memory modules fail, cpus fail, fans fail, storage drives fail. Sometimes those are correlated --- the HP SSDs that failed when the power on hours hit a limit (two separate models) are going to be pretty correlated if they were purchased new and stuck into servers at a similar time and then on 24/7. Most of those failures aren't that correlated though. Software failures would be more likely to be correlated though, of course.
The key thing is to really think about what the cost for being down is, how long is acceptable/desirable to be down, and how much you're willing to spend to hit those goals.
I can't understand this. I think transferring servers would be the the least of problems. Its the transferring of database and maintaining consistent version of databases in both the locations. Moving the snapshots after every X minutes doesn't maintain consistency. I would like to read about any company that is able to do this, as honestly it sounds really hard to me. Is there any writeup of IBM thing you mentioned?
https://news.ycombinator.com/item?id=23471698
TLDR is connectivity to and from the IBM cloud datacenters (which includes softlayer) was generally unavailable, globally, for a couple hours. If you were in multiple IBM datacenters, you were as down as if you were in only one (mostly, I was poking around when it was wrapping up, and some datacenters came back earlier than others).
> Its the transferring of database and maintaining consistent version of databases in both the locations. Moving the snapshots after every X minutes doesn't maintain consistency. I would like to read about any company that is able to do this, as honestly it sounds really hard to me
The gold standard here is two-phase commit. Of course, that subjects every transaction to delay, so people tend not to do that. The close enough version is MySQL (or other DB) replication, monitor that the replication stream is pretty current and hope not a lot is lost when a datacenter dies. There's room to fiddle with failover and reconciliation; I recommend against automatic failover for writes, because it gets really messy if you get a split brain situation --- some of your hosts see one write server available and others see another, and you may accept conflicting writes. A few minutes running like that can mean days or weeks of reconciliation, if you didn't build for reconciliation.
The main IT guy went on holiday and one of the cover guys from another office decided to tidy up. He unplugged the server and thought (and told me after his thought process) "if anyone was using it, they'll let us know".
This was the one, single box for the whole website - no one else was monitoring (even though the central office had a proper, dedicated web team) and the assumption was I was sysadmin.
An hour later I'm sprinting down the corridor to find out what the hell happened and why I can't even SSH into the box.
We put a sticker on the case saying not to unplug it after that...
Anyway, by the time I was there it was still a '70s-vintage large computer room but now massively overprovisioned on space, cooling, etc, particularly with most IT functions having moved to corporate. A decision was made to repurpose part of it as a test lab and move all the actual remaining equipment to three racks in the corner.
I'd do about two servers a day in between other things, taking advantage of redundant power supplies to transfer the PSUs one at a time to extension cords, swap to a long network cable fast enough that TCP sessions probably didn't time out, and then unrack onto a hydraulic lift card and do the same procedure the other way.
I presented this at the start as far from a guaranteed strategy - that it would minimize downtime but there would inevitably be some due to mistakes. None of this was really that critical. There were a few devices that were pretty old and poorly maintained, we agreed up front that if these lost power for some reason and then failed to boot, we would just say they'd lived long lives and purchase replacements.
I guess the point is that this whole situation was kind of unusual and I would generally not recommend doing this, we were lucky that all the equipment left had stakeholders that acknowledged it was legacy stuff and they could tolerate losing it.
The irony is, of course, that it went perfectly. So far as I know there was not a single problem experienced through the whole thing. I even managed to swap the phone lines to the (surprisingly busy!) legacy fax server when each was out of use.
And the sysadmin let that opportunity pass??
I'm just picturing the networking equivalent Indiana Jones swapping the idol with sand in Raiders Of The Lost Ark.
The client would immediately refuse to pay anything because he was very clear he wouldn't pay a thing if there is downtime.
Then, the next contractor would be super quick to judge you and the situation, reinforcing that you were an incompetent idiot and the client was right to kick you away on the spot and not pay a dime.
Glad it went well in the end. There is so much to lose for the person trying to help.
I am having trouble finding a reference to it now, but I've heard patio11 refer to this as "the Japanese no". Don't ever say "no" directly, just quote an astronomical price.
--
[1] https://www.networkworld.com/article/2220097/what-microsoft-...
And now I know that Coldfusion is absolutely miserable to code in (and the client tried to dodge their bills!).
So you quote a price that is high, but not so high as to destroy the relationship. I call it "plausibly-deniably-high".
You also have to gauge the context of the other party in the negotiation. This technique works best when you accompany the quote with some kind of description matching the personalities. Some people are swayed by a description of the additional time it takes (the billable hours mentality). Others are swayed by a description of the additional risks you are bearing on their behalf to deliver the outcome. Still others are swayed by a description of the de novo technical challenges that no one else has ever attempted before. The list goes on, and is a fascinating study into people.
This is where a real salesman (as opposed to an order-taker) earns their keep, where they know how to read a room and craft a response, messaging and after-meeting socializing that takes into account all those perspectives simultaneously from the point of view of the other party.
A friend of mine with a consulting biz was requested by IBM to handle a job in Turkey. He didn't want the gig & told them so repeatedly. He finally decided to tell them the most ridiculous price he could think of (like appending two zeros to the number). He said they didn't even flinch and he was on the plane to Turkey the next week for six months. But he did say that it was pretty much worth it in the end (but only because of the pricing).
Seriously, this sort of dynamic is why the world works as well as it does.
Yet they are not a panacea.
They suck at preventing problems related to:
* tragedy of the commons - tend to create & magnify it
* long-term disaster planning / tail risk - e.g., stockpiling resources for natural disasters, pandemics, etc.,
* preventing foolish development, e.g., on cheap land subject to flooding
* self-creating safety systems for workers, consumers, environments, etc. -left to their own devices, markets always do too little-too late
Market systems literally often need to be saved from themselves, e.g., when overfishing will literally kill an industry by driving extinct the very thing it depends upon
Price gouging is nowhere near a reliable method of disaster preparation as actual expert planning.
The stockpiles you speak of are usually just ordinary current inventory marked up by an order(s) of magnitude.
Also, stockpiling goods is not the only thing needed for disaster preparation. One must also stockpile services, i.e., have the right people recruited, trained, equipped, and ready to respond. Prime examples are military and firefighters, who spend a much time & resources training, and little time actually fighting the wars or fires.
Yes, but funding allocation is hard.
That'd be the case only if it was sudden. If entrepreneurs had the time to think and plan, they'd come up with stockpiles that they rise the sale price for when the time comes, calculating to use that future sales price increase to offset the increased bound capital and storage expenses of their large(r) inventory.
Military is a bad example, but firefighters do train a lot. But that's also due to them needing to respond within hours at best, instead of weeks/months for most wars. I'm referring to the majority/bulk of them, not the leadership hierarchy.
For what it is worth, if a customer of my previous (salaryman-heavy) employer asked for this, we'd tell them an actual no, which is extremely rare in client relationships in Japan. A contextually appropriate "no" for something which is less absurdly wasteful of engineering time to no purpose would be "That sounds difficult. We could explore options to do it, but perhaps you could accept an hour of downtime in the dead of night" then bargain down to 15 minutes.
https://www.consumerreports.org/cro/news/2014/04/record-play...
"The stylus did not jump the grooves even when the car was moving at various speeds over broken pavement, cobblestones, and deep holes."
At my last job, we had 2 airplanes with 5 computers each with 6 disks each mounted in an aircraft. These were regular servers from Dell, not special hardened or resilient hardware or anything. So 60 or so hard disks flying around. Takeoffs, landings, turbulence. Two flights per day, 3 hours each flight, 6 days per week. So 626 landings per year.
Disk failures were not particularly common.
(They also tried spinning disk machines on buses which also failed quickly but that was more the grime and electrical noise than the motion, IIRC. Then they tried mini-servers running from CF and the motion would slowly work the CF cards out of their sockets. The company did not last long.)
Networking/spanning tree loops, arp table mismatch/corruption, the switches at the destination being misconfigured are all realistic problems that would result in downtime here. The normal way you do this is with live migration from hyper-v or vmotion from ESXI. If the initial migration is not successful, you just leave the server powered on while you address the issues. Once the VM has been migrated you can do whatever you want with the original server without worrying about downtime.
A couple of months I joined, a room full of customers chewed us out for not publishing our vmotion compatibility tables. After 4 hours of chewing out - they then told us they reverse engineered the compatibility tables and reorganized their entire data center to conform to vmware vmotion. Then (of course) we worked with intel to make sure the compatibility matrix worked in the future.
I realized at that point that I joined the right company.
I mean if he had access to technical resources who were willing and capable to do this for him, he chose to do it.
Then again, could easily be wrong
When you get clients describing things like this, it's possible they've been promised things about this server before by other consultants that didn't pan out. They don't want to give you the full details because then you'll recommend a different route that they don't want to take (justifiably or not).
It's easier for them to frame the problem to a consultant in a way that allows for only one potential solution, even if perhaps better ones exist, because the guy in charge of making the decision isn't technically skilled enough to assess whether others proposed by consultants are as viable.
And, of course, one might read a little into why there exists a "boss" with such a highly-critical IT need that is hiring a consultant to do work like this, and thinks that threatening to not pay at all if there is any downtime is the best way to do it.
I mean, what if they opened the door to this closet and it grazed a power cable on the floor and the machine just shut off? Why even bother staying around to bring things back up? It wasn't your fault and there's already downtime: you're not getting paid.
I don't really understand the "ranty" tone. The client had very specific requirements and the author came up with an effective solution and was fully paid to deliver it. Sounds like a win for everyone.
I read it as "if you take small enough discrete time intervals they won't overlap with any downtime". Or in other words "no downtime between downtimes". Yes, it's very in line with your video.
For physical servers the uptime is typically quite small. Of course Google isn't optimizing for server uptime so it isn't fair to say "well even Google can't do it".
Hard work is wonderful stuff. Days and weeks of it can save you whole hours of planning.
It appears that the client didn't have the specific requirements on initial consult.
It's quite sad IMO, don't recommend to go there unless you want to have a bad day reading about the most horrific work environments and bad practices in the world.
It's not supposed to be negative per se. It's supposed to be entertaining.
It's an art form. It's not everyone's cup of tea, just like horror isn't everyone's cup of tea. But people often watch horror movies for catharsis, not because they want to be depressed and wallowing in self pity.
Storytelling is often about educating people about things you can't speak about more directly. It's often a way of sharing wisdom in an inoffensive manner and one that will stick because people will actually pay attention, unlike when you are giving them some dry lecture about some problem they haven't yet had and don't yet care about.
But if you entertain them, they will read it anyway and that story may stick with them. And then six months or a year later when they have the same problem, they will actually remember how someone else handled the same issue and it will turn a potentially nightmarish scenario into "Meh, I just did the same thing that guy on Reddit did to his shitty boss/customer/coworker. Worked like a charm. Moving on."
This is literally the basis of human interactions, thats how we humans work at every scale to form friendships/families/societies/nations.
It just stated that the specific pressures and rewards present in most reddit communities tend to encourage this specific style of writing.
The flair for the piece is "Rant." That's an official category for the sub. There are going to be expectations surrounding how you write when using a tag like that.
Just like Hacker News! Here's a clue--it's in a subreddit where these types of stories are welcome.
They're famous for not being a cheery bunch. Because reddit's demographic does swing younger the sub used to be filled with endless posts about being socially incompetent or possessing 0 business craft.
Does anyone know if it improved?
'Moving online webserver using public transport'
Luckily one employee was working from home (rare at the time!) and had a copy of the entire movie on her desktop computer. Which they very carefully moved back to the office and were able to restore from that.
https://www.reddit.com/r/uptimeporn/comments/1kf26r/moving_a...
When it's done for pay rather than for fun, and payment is conditioned on zero downtime, I hope they charged a premium to make up for the risk of no pay. Offhand, I don't know what's a good way to do that -- I've never had a consulting client demand terms like that for billed-by-the-hour work.
Risk = client risk * task risk.
Client risk is based on your past experience with the same client. If they're prone to demand last-minute changes or stupid stuff, they get charged a higher rate on every project afterward. Jacking up the client risk factor is also a nice way to fire a client you don't want.
I've found that when people are being unreasonable it is because they haven't split out their true needs from their first idea of how to meet those needs. In this case the true need is zero impact to users. The owner translated that to "zero downtime", and then didn't accept alternative solutions that still would have met his true business need.
So I listened to him whine for a couple minutes, then tossed a dollar on his desk, told him that would cover it so he could shut up now, and rebooted the server.
Warning: you should probably only try this if you are good friends with your boss. That boss had been my best friend for years before I came to work for his company.
Things were so much simpler back then.
I used to be all about long uptimes. I eventually started seeing long uptimes as a negative though. A long uptime probably means patches have not been applied.
I think the cult of runtime came about simply because it was impressive that a personal computer could stay running for more than a few days when most of the world ran Win95. And because development cycles were longer and there weren't a lot of network threats.
..It's interesting how pop-culture and your chosen profession intersect, at times.
There was downtime.
Promptly after our cluster settled into this wonderful new facility, a cooling pipe in the ceiling leaked on it, frying 1/3 of our nodes.
Now that I am older, I don't think I would do it anymore, too much stress for a small reward. Also today, most of the time, I am able to "talk out" customers of crazy requirements, while I would just have said "OK let's do it" in my younger years.
(Not to suggest it’s bad, just different now that a primary assumption about people work in the office is less true)
This article provides an example of how when you operate on prem, literally any crazy option remains on the table for you. If you asked your cloud provider to do this, it'd be a no.
That was my main take away from this. Endeavor to be the sort of person who can refuse clients, the entire idea that "the customer is always right" enables so much ridiculous behavior.
The solution I came up with was to fashion a custom male<->male power cord, like a gender changer, from some broken ATX PSU scraps we had laying around. By rearranging the power sockets from multiple donors, two male power cords could be connected on a single enclosure. Internally the sockets were simply bridged, otherwise the PSU was basically gutted.
With this goofy metal box having two male power cords dangling from it in hand, I just used a very long extension cord plugged into an outlet on the same AC phase as the existing server's power source. The extension cord powered one of the bridge cords. The other bridge cord plugged into the server's existing - and hot - power strip, forming a redundant power source. Now the power strip could be unplugged from the primary power source without losing power, and we just moved the server to the new location with the bridge box and power strip in tow.
If memory serves the only tricky part was determining which outlet at the new home was on a compatible circuit. We didn't have much in the way of electronics tools, no oscilloscopes or anything. Even the soldering involved to make the bridge box was done using my personal soldering iron, which just happened to be in the office because some of us raced RC cars there after hours.
I think I just used an incandescent desk lamp to verify a normal brightness on the bridged circuit before proceeding with the server, but it's been a while.
I wonder how many people have fashioned AC power cord gender changers throughout history... :)
You could clone the VM to another instance and record commands going to VM1 and replay them to VM2 after 5 minutes.
This whole brain fart of mine doesn't make much sense but if you play along with it, does it still count as a downtime or just very high latency?
If you do a ping during “bad weather”, you can see that they buffer up to 5 minutes of packets (i.e. there will be no communication for some time, then you’ll receive a bunch of them with a huge latency with sequences intact).
So I would assume a lot of software could even work that way. I think a lot of software don’t set any (TCP) timeouts at all.
5 minutes of unplanned downtime in a pub/sub setup could easily go unnoticed, since that setup is typically tuned for long timeouts and/or repeated retries.
That sounds like I'm being snarky but I mean it - whether an actual legal contract or just the documentation given to users, any system where downtime matters should have some discussion of what impacts downtime can have and how it's measured and managed.
That documentation is what defines "downtime".
I'll add that what you've described is a sort of low-fi manual version of DB replication (https://en.m.wikipedia.org/wiki/Replication_(computing)).
It was already plugged into a UPS, but they had to cut one of the posts off the rack to get the server out without unplugging it, then they plugged that UPS into a bigger UPS on a cart and wheeled it to the new data center they built out in the building next door.
The world was much different at the time -- this coloc provider had a good reputation, yet.... they had a keg of beer in the corner of the server room and a stack of adult magazines in the men's room.
Perhaps they should have just told the customer they couldn't find it: https://www.theregister.com/2001/04/12/missing_novell_server...
The other way I'd do it is more similar to described. Create redundant network paths to the server, then cut one.
Like finding people who argue against revision control systems, it's really quite a challenge convincing people why things like this are a bad idea - after all "it works!".
After several attempts to understand the binary format I gave up and ended up printing tabular reports to LPT1 which I connected my laptop to, extracting it and rebuilding CSV files.
Lucky enough, printing those days were the most important feature of a business app.
But have known some large companies who have in their history, done things like this and other creative solutions to impossible problems.
I would probably still charge a much higher rate since the owner was an arse, but at least you would get back your 7-8 hours.
Live migration of VMs would have been a better option, which was brought up in the reddit comments and dismissed because HyperV live migration is spotty. While I'd have to agree with that assessment, it isn't so spotty that what they actually did was less risky.
In a way, this is Darwinism for the IT industry and I'm happy the people involved got paid well. Due probably paid as much as it woulda costed him a new server. I bet he'll never forget this lesson.
Moving a running server about 7km through public transport without downtime.
Saturday 3AM shift with a 5 minute downtime would work just as wwll. Unless this server has had historically 100% uptime this would go unnoticed.
Besides, as pretty much everyone has noted, running a zero-downtime system on a single physical machine in what sounds like is just a normal cable room is kind of nuts. Those 10h would have been much better spent to move that puppy to someone else's data center and get some redundancy.
Although reading between the lines, maybe the lease was up and they were waiting to the last minute to move it.
(means extra billable hours for the extra manhours needed to hold the umbrellas)
what a crazy world.
Naturally, I did my best to explain the laws of physics to him, but he wouldn’t hear it. In a spectacular display of Stockholm syndrome I did my best to appease him for four years, but, as many of you can surely predict by this point in the story, I failed in every possible way and eventually gave up. Just wish I could have my four years back.
I was glad to read that OP at least got paid well for his efforts.
I usually get fired from such positions in less than two.