“A Windows 7 deployment image was accidently sent to all Windows machines”
it.emory.edu
it.emory.edu
"To make error is human. To propagate error to all server in automatic way is #devops."
Frankly, I'm surprised things like this don't happen more often. Kudos for the incident management. Also a big plus for having working backups, it seems.
"In devops we have best minds of generation are deal with flaky VPN client."
"Single point of failure in private cloud is of usually Unix guy with neckbeard."
These are gold.
Edit: based on the above advice I once grew out a neckbeard while going through a multi-month rollout of a large product. It itched like crazy, but I did work much faster to get rid of it.
So true. I'm on the receiving side of this..."No you can't work on that multi million deadline project of yours...the only way to fix the VPN is to re-image the machine back at head office [an international flight away]". Me..."Could you repeat that?" And thats a Cisco Enterprise VPN...(turns out IT was right...re-image & avoid conflicting software is the only solution). So much for Cisco...
After an hour or so of troubleshooting it's usually better to go with the reimaging, since all you / the user wants is to get back to working.
Ideally I try to get the entire broken machine captured and the user issued a new, fixed machine because then a fix can be developed and documented, but for those who end up in a new failure mode, it sucks. And with something like the Cisco VPN Agent? That's not uncommon at all...
Definitely. In our case its 8 hours minimum though for a re-image. Somehow the FDE makes pulling the old data off the machine slow.
You've got my sympathies though - I'd not like to be the one doing the IT in these cases. Can't be fun troubleshooting IT with that kind of time pressure.
If this is something that smells of a bigger problem (or has been seen elsewhere) then I push for them to get the user a wholly new machine, capturing the old one for analysis. If the user is given an upgraded machine, then there is usually little resistance, even with the downtime that'll be incurred.
On the upside, if the issue can be reproduced readily, from this we can almost always get root cause and put a systemic fix in place. If it's sporadic... Well... I'm sure you understand how it goes trying to fix something that you can't yet reproduce. ;)
(I'd love to troubleshoot your slow data backup issue... That's the stuff I rather enjoy.)
I'm not directly involved with the tech side so I don't know the details. I gather they pull the old data off the disk using some offline low-level tool though (like you would for harddrive damage recovery). Between that and the encryption its somehow very slow. No idea why its like that though.
>get the user a wholly new machine
I wish it was the same here. They just give loan machines :/
Couple of reasons. Each country rolls their own custom image. Plus I need an office that has the encryption keys for the full disk encryption. Plus only 3 offices globally carry copies of my data (used when they can't pull the data off the hdd).
If I'm flying anyway I might as well go to home office - I know they have all the required stuff for my laptop.
"Turtles all the way down" is a "a jocular expression of the infinite regress problem in cosmology posed by the "unmoved mover" paradox."[1]
But not English, I can't make sense of them.
They do. This happened to the largest bank in Australia mid 2012[1]. Very similar circumstances. I've been told that SCCM's UI doesn't help here- something about the default action when nothing is selected to apply it to all devices managed by SCCM. Someone more familiar with SCCM may want to correct me here.
[1] http://delimiter.com.au/2012/07/30/disastrous-patch-cripples...
I once worked as an admin on Solaris boxes at a big pharma company. There were ~77,000 users in their LDAP directory. I was very careful.
This was a small college, so the IT guys just went round explaining and told us to log out and log in again. Applications re-installed. No data loss and so no shouting.
Stuff happens. We did 'assignment action planning' that morning: mind maps, essay plans and research ideas. Results better than normal anyway
First are the people who believe in the mission. They are really good, and are willing to take a cut in pay for some combination of social good, great working environment, etc. These kinds of people tend to be forthright about problems.
The second group are people who would struggle with the demands of the normal corporate world. They are getting paid less in higher ed and are worth what they are getting paid.
Furthermore, the profs were always happy to see me coming because I fixed their broken stuff without pointing any fingers.
It was heaven.
Long story short, the sysadmin was hired back and paid more than most of the profs. Academia may tend to skimp on salaries for certain positions, but sysadmins probably shouldn't be one of them.
Disclaimer: dev at a university.
That said, if you get a cushy job, it can stay cushy for a long time.
Note: I am not saying all university support staff are like this. Some definitely are though, and they're probably the reason why good people sometimes find it hard to be properly remunerated in academia.
Certainly not everyone is like him - but I'd wager every university has at least a couple of people like him (we definitely had one, again, in physics)
It was both direct and funny enough that I was only mildly annoyed that the cluster was down.
Wow. That is very unfortunate, to say the least...
Not in a "haha, what a bunch of morons. Serves those jerks right!" kind of way, but more in a "oh dear, that's the worst thing that can possibly happen! Oh no it gets worse??". I've been through IT catastrophes (and caused a couple myself) and I could easily see this happening to me. Still, it's funny as anything.
This was amazing reading. Reading such a detailed wrap up of an IT team going through my worst possible nightmare was enlightening.
One of the programmers decided to eliminate some code by combining the two functions, with a switch to control whether /LOAD or /EDIT was used for the first statement.
There was a bug in the program, and the edits were sent down as loads.
A guy I knew, Barry, was the main operator that night. He started getting calls from the stores after around 10 of them had been reloaded with 5 or 6 items.
Barry said it was the first time he got to meet the president of the company that day.
Since a reformat was done to the affected machines, does this mean that researchers' datasets, drafts of papers, and other IP were lost? Or were researchers' machines not affected?
If the SCCM server was pushed the "update" then there doesn't seem much hope for other machines? Surely no rule should be able to format the server running the ruleset; seems like a failsafe failure there at least.
Lawn care error kills most of Ohio college's grass
http://www.wral.com/lawn-care-error-kills-most-of-ohio-colle...
Just to clarify for the passersby, it isn't the University of Ohio, rather, Findlay University
> Findlay University did not release the name of the company that made the mistake but is working with the business' insurance company to pay for it to be fixed.
Findlay apparently doesn't value transparency as much as Emory.
> The university says grass was killed on as many as 54 of the campus' 72 acres.
One of the lessons from my college days (informally acquired, take with appropriate quantities of salt) was that walking barefoot was as much a risk for chemical exposure as puncture wounds.
How on earth do you "accidentally" load up enough weed killer to treat 54 acres of grass and never realize its the wrong stuff?
Bad stuff doesn't just happen to Windows networks.
* As soon as the accident was discovered, the SCCM server was powered off – however, by that time, the SCCM server itself had been repartitioned and reformatted.
http://channel9.msdn.com/Events/TechEd/NorthAmerica/2014/PCI...
I guess that's how robot apocalypse is gonna look like.
What is the Latin for office automation?
It's similar to a database firehose: If you accidentally start deleting all data you should have a quick working backup ready to quickly bring the dead box up to production.
> As soon as the accident was discovered, the SCCM server was powered off – however, by that time, the SCCM server itself had been repartitioned and reformatted.
That made me laugh. Poor SCCM server :)Unicast fail.
However, it looks like they handled this accident the best they could! Perhaps this accident would not have happened at a more reliable IT department.
So far all the signs have indicated they are doing great in recovering. I just hope there won't be onerous processes and restriction afterward due to desire on "make sure it won't happen again" stance.
my roommate works at the emory library and has had a fun slow week there of coming home early many days because no one could do work. they were apparently also given laptops as an interim solution, but those somehow also wiped themselves eventually (?).
poor IT people...just as they're starting to get a handle on the actual sitation it starts blowing up on the internet.
In my days of university tech support.
It "mildly" amused the *nix operations guys to see all the "point and click" colleagues panic.
http://www.youtube.com/watch?v=jR6xbulUmsg
"Yay, cloud!"
https://news.ycombinator.com/item?id=7758203 is a dupe of https://news.ycombinator.com/item?id=7758193
That is a real principle in interface design - if something would be really, really bad to activate unintentionally, make it really, really hard to activate.
If you design a nuclear missile facility, you don't put the "launch nukes" button right next to "check email" and "open facebook".
Same way it shouldn't be easy for users to delete or corrupt their data by accident due to some omnipotent action innocently shoved right in between other trivial actions.
I wouldn't blame the person who triggered this re-imaging process. I'd blame those who designed the re-imaging interface, to allow it to happen so easily by accident.
The system should be able to assess the scope of a task, and ask you to confirm 10 times if it has to, in blinking red dialogs, to make sure you really want to do what you are doing.
Of course, it's crucial that "clicking 10 times" is not the default behavior for any trivial action. Or boredom and the subsequently formed mechanical 10-click habit of the operator will kill the effectiveness of this approach...