Computer Crash Wipes Out Years of Air Force Investigation Records
govexec.com
govexec.com
A really easy way to fix this is to make a law so that if government data such as this gets lost, then the cabinet level person responsible for overseeing the agency immediately loses their cabinet level position and is barred from further work with the federal government.
This would do wonders to align incentives all through the federal agencies when it comes to data safeguards.
Edit: it also inspired a variation on an old joke about how the same word means different things to different people. For example, suppose the order goes out to "secure the data center."
Marines report back, "We have destroyed the data center."
Army reports, "We have killed everyone in the data center and are holding the position."
Navy: "We locked the doors when we left for the day."
Air Force: "We signed a three-year contract with an outsourcing company, with an option to extend at the same price for ten more years."
Which of course means that just because I don't think Hanlon's razor applies doesn't mean I think they're competent.
I've had this thought before. Can you think of a general, defensible clause to add to Hanlon's razor to account for this?
You need to treat it all as malice.
Unless that malice is aimed at keeping an individual out of trouble rather than harming others.
"Never attribute to malice that which can be adequately explained by stupidity. But don't rule out malice."
P.S. I think legislating accountability is an issue the public needs to debate. In all countries. The US has seen numerous gross negligence/errors (heck one is too many) in various levels of govt (fed, state, local). I hope we can come up with a better solution than 'fire them and bar them' and even mine of 'demote/retrain'. There has to be a better way to legislate govt accountability.
It's local data corruption here. At worst you loose the data since the last backup.
Now, in the army, they have redundancy procedures for everything, and you want to make us believe the one server used to keep them in check, not only had crash beyong repair, but has no backup ?
Many of the things that cause data loss are just simple mistakes caused by a failure to review changes carefully, even when you have multiple levels of review- I would dare to say especially when you have high confidence that someone else is reviewing your changes.
"Manual" data changes, e.g. executing SQL statements or scripts that execute SQL that aren't a part of your application, in my experience are the most common cause of data loss.
After manual data changes the second most common in my experience is not understanding what you are doing. For example, you might take a chance on an upgrade that fails because you have to meet a deadline.
Changes to application code are next. Typically when making changes to an application, a little more thought may be put into it than a one-off data migration or change, but if you are under time pressure, don't know what you are doing, or are assuming someone else will catch your mistakes, you could easily screw everything.
Following this- mistakes that cause hardware or software failure. I worked at one large organization where storage arrays with various power backups were just "turned off" by a contractor that didn't understand the impact of what he or she was doing.
Finally, you might have configuration issues or the hardware might just fail.
Really, there is no substitute for having your data backed up frequently, and in a way you know how to easily restore and have tested, by building another machine from the ground up to replace it and documenting and practicing that well. Very few do this frequently. And even if you do- what if all of your hardware were destroyed? Can you easily go out and buy something off the shelf with instructions you have in your head or stored safely around the world and rebuild everything?
And we are not talking about "any data". This was cleary very sensitive data they new they needed to protect.
Either the Air Force is failling at being the very thing it's been created to be (which I doubt) or something is fishy (ocaml razor).
This is "We had a crash and lost over a decade's worth of data" - yeah, that shouldn't happen. Ever. If you don't have a working backup from sometime in the past 13 years, -you are doing something wrong-. That's not a small series of failures or mistakes. That's either staggering levels of incompetence, or maliciousness.
This is decade of data. It should have been on in cold storage, on tested media, in multiple locations. No single error, or small chain of errors, should have enabled this to happen.
Ah, I know it's the way it is in some organisations, sure. But is it a good way to handle people screwing up? People always screw up. If you start firing them for screwing up you end up with a pool of inexperienced workers who keep screwing up in the same way. If you keep them around, they learn and never screw up in the same way again.
then you have ignorants, that in 2016 will not even think that mission-critical system should have SOME backup, off-site. people that will put amazing amount of energy into little political games, backstabbing, or just plain old incompetent lazy ones. those should become exemplary cases.
investigation should sort out which is which, but most will fall on either side of spectrum. usually it's really not that hard.
Incumbent trusts local IT manager and local IT manager has contract with external company.
The question becomes: "Did the contract with Lockheed Martin cover any form of regular backup and additionally some form of long term open format archiving?"
Perhaps some checklists or standards that all long term record keeping systems need to comply with would be good.
No, I'd slap prison sentences on everyone involved, not just loss of government employment.
I see a fantastic way to completely abuse this. Need to get someone who disagrees with you kicked out? Hire some corporate espionage people to take care of that work for you!
I was installing some service packs at about 5pm. Windows NT4. 5pm, everyone's gone home, safe, right? Server asks if I want to restart, I say yes. Then a guy walks into the server room, asks if I had restarted the server. I say yes. Whoops, OK, I guess someone was working.
Next day, the database was dead. Seemingly because I had restarted the server while someone was using it. It was a custom database job, and the contractor who had made it was on vacation in some other country. Here's the kicker. Usually tape backups have rolling tapes, right? But we hadn't started rolling tapes yet. So the good backup from 2 days prior was overwritten the night before with the crashed database. Ugh. And I was due to go back to school in 2 weeks in another city. Ouch. I still feel bad about that to this day. No idea how my boss ended up.
Never ever let someone who doesn't know what they're doing mess around with your mission critical systems where the consequences are not recoverable. Always guide them. Although in my case, it seems I may have known more than my boss at the time, so maybe wasn't possible for me....
And the answer when shit happens is "outsource! Go to the cloud!" -- so that local managers won't be responsible when the snafu happens and data is lost (if it happens to Google, it can and will happen to anyone).
Progress, eh.
No, back in the day, you could still wipe years of records because most people only used a single filing cabinet with no duplication or off-site copies. "Destroyed in a fire" was the canonical way to destroy all records, easily. This story has nothing to do with the introduction of IT. Incompetence is still incompetence.
Bernard: Shall I file it?
Hacker: Shall you file it? Shred it!
Bernard: Shred it?
Hacker: No one must ever be able to find it again!
Bernard: In that case, Minister, I think it's best I file it.
Yes, Minister - The Death Listhttp://www.archives.gov/st-louis/military-personnel/fire-197...
We know a lot more about conquered Mesopotamian kingdoms than we do about successful ones, because step one after taking over a city was apparently burning it down, or at least burning down the palace.
The only data loss to Google that I was able to find was [0], where a single datacenter was hit by lightning 4 times in short succession, and 0.000001% of disk space in said datacenter was lost. And only because the lost data consisted of recent writes that had yet to be backed up. And only because the lightning affected the AUX power systems in a way that they have since committed to fixing (or perhaps have fixed).
IMO, cloud storage is the most reliable way to store data, and Google the most reliable. I never know if I'm going to drop my laptop and lose all work. Or if the harddrive in my server is going to give up and burst into flames. But the great thing about cloud storage is that it abstracts away hardware components and failures so it's somebody else's problem. And they worry a lot more about data loss than I do. Cloud data storage is designed with the assumption that hardware is faulty. Spread redundant copies of data to geographically isolated datacenters, and you're much safer than some admin of averagejoe.com keeping his data on a single harddrive in his closet.
[0] - https://status.cloud.google.com/incident/compute/15056#57195...
Does Google publish any durability numbers? Amazon claims 99.999999999% (11 9's) for S3 and Glacier (and considerably worse for EBS volumes). How does Google compare?
(disclaimer: my life would probably be ruined by losing the gmail account I've held since private-beta days, because life is short and I'm shit at sysadmin; but there's a big difference between a lone old geek and large organizations...)
[1] http://www.dailymail.co.uk/sciencetech/article-2548010/Has-G...
[2] http://www.datacenterknowledge.com/archives/2011/03/01/googl...
For more serious data, I have to imagine that SLAs generally define data retention periods even after exiting the agreement, right? (I legitimately don't know)
Obviously it doesn't help if the company goes under, but it's not as simple as the parent commenters statement.
That said, over the years I've witnessed all manner of billing and payment fuckups. I'm not trying to say that cloud is bad -- but that billing snafu is a risk factor that you need to understand and account for.
Prime time for burying bad news.
Or something along those lines.
This phrase made me laugh, but I don't really understand what it means, can you explain please?
It made it sound like the staff were running around and being very energetic in their operation of the paper shredder.
Much of the data is backed up, and more of it will be retrievable from original sources. Maybe not all of it; I'll be very interested to see what gaps remain when they declare the issue resolved.
The most common move is to cut the highest paid section of the workforce after they acquire the service contract from a competitor. All major defense firms do this, which leads to problems maintaining institutional continuity.
>> It’s possible that some data is backed up at local bases where investigations originated.
Very convenient, that flood.
James Hacker: How am I going to explain the missing documents to "The Mail"?
Sir Humphrey Appleby: Well, this is what we normally do in circumstances like these.
James Hacker: [reads memo] This file contains the complete set of papers, except for a number of secret documents, a few others which are part of still active files, some correspondence lost in the floods of 1967...
James Hacker: Was 1967 a particularly bad winter?
Sir Humphrey Appleby: No, a marvellous winter. We lost no end of embarrassing files.
James Hacker: [reads] Some records which went astray in the move to London and others when the War Office was incorporated in the Ministry of Defence, and the normal withdrawal of papers whose publication could give grounds for an action for libel or breach of confidence or cause embarrassment to friendly governments.
James Hacker: That's pretty comprehensive. How many does that normally leave for them to look at?
James Hacker: How many does it actually leave? About a hundred?... Fifty?... Ten?... Five?... Four?... Three?... Two?... One?... Zero?
Sir Humphrey Appleby: Yes, Minister.
(Source: http://www.imdb.com/character/ch0030014/quotes)
Edit: Although apparently a similar real-life incident occurred when Hurricane Sandy wiped out a significant portion of an FBI record archive: https://nsarchive.wordpress.com/2014/09/16/archival-neglect-...
[...]
>"We've opened an investigation to try to find out what’s going on, but right now, we just don’t know," Stefanek said.
I wonder where they're putting the files for that investigation.
I wonder if their backups were corrupted too, or if they just overwrite the old ones with newer ones as a cost-saving measure. Or whether no one ever tried restoring from them and for whatever reason they never could have been restored from.
Since people rarely prioritize it until it's too late, maybe 'restore from backup' drills should be a common thing.
Ours (at a state university) deals with petabytes of data. We've had to escalate to IBM's Tier 3 support a couple of times over the past decade, but we have never lost a file.
I asked the one of the contractor ITs when was the last backup. Response: "Do you want me to do a backup of your workstation tonight?"
WTF!!! I am no IT slob but the minimum is Mon-Thur backup of new and changed files that day, and Friday night was full network backup.
Why wasn't this done?!!!! The tape backup makes too much noise. I guess the bitch couldn't read her endless Romance Novels because of the noise.
I find an intern, not a full employee, working on a design file in the account of another engineer. WTF!!! We believe that since the intern was leaving soon, it would appear that the design was the work of the engineer.
Security violation!!!!! The contractor IT people swore they did not know this was going on...... BS!!! With so few on the network at any one time, and in the mornings only the intern was on the network, these contractor IT people were unaware!!!!
The government IT person knew, but she did not want to report it because she was well aware of the contractor's history of retaliation.
BTW..... during the investigation, the government engineer asked where was it written that they could not share passwords!!!!! Can you hear Snowden laughing at this.
All swept under the lumpy rug....... the violations would make the organization look bad.
As for me..... I got the IR treatment.... that's Isolation and Retaliation... then after 2 and a half days, supervisor sent me an email as to my where abouts.... my reply: Retired as of 4:00 PM PST, 3 days previous.