Devops Horror Stories
statuspage.io
statuspage.io
- Somebody deployed new features on a Friday at 5pm.
- Fifteen hundred machines running mod_perl.
- Supporting Oracle - TWICE.
- It turns out your entire infrastructure is dependent on a single 8U Sun Solaris machine from 15 years ago, and nobody knows where it is.
- Troubleshooting a bug in a site, view source.... and see SQL in the JS.
Now we deploy dozens of times a day and I never get called on a Friday night because someone did something stupid.
Edit: I do get called when I did something stupid and it broke the deployment system. But that's gotten much rarer lately.
Behind the scenes it was a decentralized continuous delivery system. Very cool stuff, highly automated. Reduced a lot of work and sped up development cycles from months to minutes. Served quite a large software development organization (1000+). I think we had 1000 servers in 5 datacenters around the globe.
Nowdays I'm working on an open source version of that system, it's still missing few critical features but hopefully I'll get the first release out next spring.
btw, I'm looking for projects/contacts that would be interested in trying out how the system would fit their needs.
Do you have any sort of github/project page?
Cool, sysadmin at a research university sounds like a nice position to be at.
Yes, it's already on github. Unfortunately, since it's missing those critical features it's not easy to see how the whole system is going to work. If you to talk just drop me an email at gmail. mikko.apo is the account.
Personally, I think it's better to release at 5pm on a Friday. Once people stay late a few times to fix their broken shit they'll be smarter about not checking in crap.
Even if it only happens once in your career, once you've had a dev push out code at 5pm friday night, jet out the door and hit the bar, meanwhile you (the sysadmin/ops on call) get woken up at 1am by site down alert, and have to debug/rollback the changes while the dev who pushed them is unreachable, you learn to really avoid friday evening pushes. Fool me once...
Its sad that most places dont have a proper technical copy (with a full copy of live data) to do full tests on TDD is all very well but you need to test the entire system.
Or when the bug is only triggered in specific user profiles.
Or when all the devs went on a retreat in the mountains with no cell service.
Or when a dev makes a mistake (which we know never happens to even the best devs)
Or when the only developer that knows which one of the 1000 changes that were pushed could be the one breaking, turned his phone off.
Or when a flaw is discovered in the process for the first time (which we know never happens because everyone's process is perfect, until it isn't)
Or how change management's requirement that the fix be tested and verified by all affected teams might have people staying a few hours after 5pm on a Friday when they just want to get their weekend started.
Or how 10 different people from 10 different teams might need to be called and kept to work until 2am because the change can't be pulled because the database was already modified and the old client data is already expired from cache and a refresh would destroy the frontend servers.
Or another reason.
It drives me nuts when people tell me off for saying 'yeah yeah, no, automating our entire infrastructure of 5 servers isn't really worth it right now', like I'm some unprofessional bozo.
I pretty much have experience with all but one or two of your suggested scenarios, and by now I have no patience for annoying software developers who think that using chef or puppet somehow sufficiently embiggens them to run ops on their own (of course dev ops is almost a political assault on existsing ops guys, not merely a nice new solution to existing problems).
Sigh. This is why I don't work on teams these days (if I can help it).
EDIT: Though I also agree with the sub-parent, that deploying on 5pm is fine in certain teams and certain projects, the most important thing is are the guys pushing to do the deploy going to own the deployment? Are they going to hang around for another 60 minutes to check everything is OK? Are they going to be available at 10pm or on Saturday if something goes wrong and are they going to own it? If the answer is no, then nope, don't do it.
No matter how great your processes and your code are, no test can catch everything that can go wrong in a live environment, and doubly not if your system interfaces with anything third party.
But the real reason we deploy weekday mornings is so everyone is on deck and we can get outside help if required. When I was doing system integration, the problem was never in my code, it was the vendor's. Testing can only get you so close to the real world.
this is why I refuse to do "View Source" on the HealthCare.gov website. I'm afraid of what I might see.
if ('en' === 'en') {
$('#desktop-nav .middle').append('<a class="liteac-login topnav myprofile" href="/marketplace/auth/global/en_US/myProfile#landingPage">My Profile</a>');
$('.mobile-nav-right').append('<a class="mobile-right-bottom liteac-login" href="/marketplace/auth/global/en_US/myProfile#landingPage">My Profile</a>');
} else {
$('#desktop-nav .middle').append('<a class="liteac-login topnav myprofile" href="/marketplace/auth/global/es_MX/myProfile#landingPage">My Profile</a>');
$('.mobile-nav-right').append('<a class="mobile-right-bottom liteac-login" href="/marketplace/auth/global/es_MX/myProfile#landingPage">My Profile</a>');
}because, you know, static HTML and some CRUD/lookup logic behind-the-scenes is just that hard
How did you locate it? Measuring ping latency from other machines?
Forgot about tmpwatch, a default entry in the RHEL cron table to clear out old temp files.
4AM the next morning, recursive deletion on anything wiuth a change time older than n days.
Edit: Reading tmpwatch.c: "Try hard not to go onto a different device". Perhaps old bug?
Why no, I've NEVER accidentally deleted whole file systems, I have completely earned superiority here.
Delete /proc and /dev on a running server. Thankfully not really disastrous but damn if people don't notice right away.
Thanks for the tmpwatch info btw.
Obviously that was incorrect, but the reasoning was, I think, sound.
Should be bullet proof?
Alas, the 'write to tape' scripts I'd inherited didn't warn if they couldn't load a tape into the drive.
There was a tape jammed in the drive, so the tape robot was refusing to load any new tapes, but kept on writing and restoring from the same tape over and over again.
Stupidly, we didn't do any 'check a tape from 3 weeks ago' for a while.
Lost quite a bit of data. We still have the md5sums though... Still gets shivers thinking about it
I, being mainly a software guy, didn't consider the hardware robot as something that might fail w/o error.
Now, the process looks like:
Check drive is empty. Load Tape. Write Tape Unload. Load tape "42" from anther slot. Write 'slartibartfast' to that tape. Unload. Load original tape. Restore & Compare. Unload. Load tape "42". Restore, and make sure all it has is 'slartibartfast'.
This seems to me to have removed most of the possible silent-failure situations. If anyone can think of part of this algorithm that might fail, let me know!
Apparently water kept flowing into the humidifier tray of the chiller, and the mechanical auto-shutoff never triggered. The pump didn't remove water from the tray because the power was off.
Facilities "fixed" the humidifier, but it still happened again when that circuit was cut off for work elsewhere in the building. No one caught the water overflow, and it flowed out down the conduits to the first floor. So we had flooding on 2 different floors from a single chiller.
When you do something like this you need to make sure that you have a good distribution of keys across the name space, and you need to think twice before you decide to write code to list the entire bucket. In most use cases at this scale, metadata and indexing are handled by something other than S3.
You get what you pay for.
Support is one of things that you can get along without... until you need it. Then you really, really wish you had it.
Entirely different problem from a provider that loses one internet connection and their other links can't keep up with traffic, but you can still have major problems even if you're spending thousands of dollars a month compared to hundreds.
At a previous job we hosted with {HAL} out of Atlanta. A NOC operator there saw/heard/smelled something that indicated to him that he should hit the Big Red Switch. So he did. This removed power to every machine in that part of the DC.
After management confirmed that there was no life-threatening emergency, they started bringing everything back up. Only to have machines start going down again 20 minutes later, as their local UPSes ran out of juice. Someone had to walk around to every cage and recycle them all manually.
The idea of using DC hosting providers is because the uptime, environment and usually strong network connectivity. Given the poor run they had I think the developer would have had a better uptime if he had hosted this at home on a server.
Turned out the two "redundant" providers of fiber both had fiber going through that junction...
It's 3AM. You're being paged with a high latency alert in one datacenter. You run one command to drain traffic out of that datacenter. The latency graph starts looking normal again. You go back to bed at 3:05. You look at the logs and figure out what went wrong tomorrow morning.
http://www.informationweek.com/server-54-where-are-you/65055...
- Someone on the hardware team deleted several VMs that were being used as build machines, there were no backups. That wasted around 2 days to get things back to normal.
- During a show I volunteer for: a scissor lift drove over a just run (several hundred foot) ethernet line and severed it, they had to run a new line.
- PCs running windows being set up, as point-of-sale systems, to run with static IPs on the internet, without a firewall running. Disavowed all responsibility and left them on their own for that. They would have run them unpatched too without intervention.
- Someone checked a private key into the repository. Plan of action: obliterate from all branches everywhere, delete from all build drops (which contain source listings too), track down all build drop backups on tape and restore-delete-then-recreate them. Luckily I handed that job off to someone else.
Another fun one was coming in one morning, and cleaning up after somebody used some foul PHP provisioning scripts on a customer system and had the unfortunate idea to use a function called "archive". Turned out the function didn't so much "archive" as "delete". Henceforth deletion, especially unintended deletion, was known as "shotgun archival".
https://groups.google.com/forum/#!forum/alt.sysadmin.recover...
Note that space? I didn't.
"Our distributed application produces the same type of error after the same period of time in totally different data centers. We have no idea why, but moving data centers seems to help, so we just keep doing it. #YOLO"
"We've built a product on a data store and library we don't understand even the highest-level constraints of. That ignorance bit us in the ass at peak load. We patched over the problem and continue gleefully into the future. #YOLO"
These stories should be embarrassing, but they're seemingly being celebrated, or at least laughed about. Am I off base?
calling them "terrible practices" is redundant, all devops horror stories can be characterized as exposing terrible practices if you're simply looking at the post-hoc view. it's a feature, not a bug, to make light of them. they're laughed about, but with the intent that they're not made again.
"No, wait, don't…!"
<site down>
But that's because it was properly configured so a reboot was smooth and didn't have any snags or affect other systems once back online. At another data center across the hall, if their main server needed to be rebooted (not accidentally!), it was 3 days of troubleshooting to get it back up. I learned that after the boss hired one of their admins - not surprisingly, a big mistake.
(2nd worst) delete from table where created < 'old_date' WITHOUT an account (thank god again for backups)
Lesson learned, always backup and write the WHERE clause first
This has saved me from paying the price for that particular class of mistake on a number of occasions.
Had a friend who recently took down his nic over ssh; he claimed he managed to get back in using some sort of serial over lan magic but I suspect he really just got someone on the other end to help.
Extremely professional.
Reality: Monday Morning: T1 did not get installed. Tuesday: Emergency ISDN solution (stolen from Chiropractors next door) Wednesday: Modem rack catches fire Thursday: TV ad goes Live Friday: T1 goes live. Champagne.
I have worked at several places where salespeople have sold a feature without even asking if it was POSSIBLE, much less created/deployed/tested. "We just sold [Feature X], we told them it'd be ready by [date pulled out of thin air]."
# kill % 1
Instead of %1. So I zapped the initializer, parent of everything else, logging everyone out without warning.
I had more than enough capital to avoid anything more than the deserved ribbing, but it was my Crowning Moment of Awesome devop lossage; harsh but minor screwups in the decade previous had trained me to be very careful.
I've avoided being handed the horrors of many other posters by primarily being a programmer. You full timers earn my respect.
ADDED: Ah, one big consequential goof, related to my not being a full time sysadmin but knowing more than anyone else in my startup. Buying a Cheswick and Bellovin style Gauntlet Firewall from TIS ... not realizing they'd just been bought by Network Associates, who promptly fired anyone who knew anything about supporting that product.... (At that time I didn't even know about iptable's predecessor, although given it was a Microsoft shop....)
I was fired from that job in part because I was the least worst sysadmin in the company, totally consumed with a big programming and database migration effort (Microsoft Jet -> DB2 -> DB2 on a real server), and gave opinions that others sometimes accepted and implemented without due diligence. E.g. I said "this is a competent ISP", not "you should also use their brand new email system" (which I didn't even know existed) ... visibility all the way up to the CEO is of course not always good....