The Grasshopper Outage: Co-Founders Response
grasshopper.com
grasshopper.com
I'd disagree. If you have same model, same production batch of harddrives in a RAID, the chances of simultaneous failure is elevated. Depending on RAID level and workload, it's quite possible two or more drives are loaded (accessed) in exactly the same way. If, again, they are from the same production batch, chances are they have similar production defects in mechanics or semiconductors, thus it makes sense they'd fail almost simultaneously.
Usually running raid10 for performance reasons but when I don't need fast writes I use raid6 instead of raid5 for anything beyond 8 drives.
[1]: http://techblog.netflix.com/2010/12/5-lessons-weve-learned-u... [2]: http://www.codinghorror.com/blog/2011/04/working-with-the-ch...
Here are some questions: - How many spare drives were in the system when the first and second drive failed? Netapp does not shut down volumes of storage if spares are in the system to take over for the failed drives. - How long was it really before those drives were replaced with a spare that could take over and rebuild? - Why don't you publish your DR plan and explain exactly where it didn't go as tested and planned? Point this out to show where the issue occurred that had been previously tested and shown to work properly.
I am not an employee of NetApp and not even a customer of NetApp. I have used it in the past and like all technologies it requires the care and feeding that is well documented in manuals they provide. And it requires good administrators to do their jobs to test and monitor things and ensure the resources are available for the system to work as designed.
Once you have heard both sides of the story the only thing you learn is that there are more the 2 sides to the story.
When these types of things happen the folks closest to them always leave out details or cannibalize the story so that they can be found blameless. I have been in the IT field for 15 years and seen it time and time again (with many technologies).
I read an email from Grasshopper this AM detailing to their customers what this issue was. It was so vague and left so much to interpretation that it really came across as whiny and misinformed. It was extremely unprofessional to apologize and then blame (without full explanation or root cause).
There's more to this than meets the eye. Believe me.
Since you all are replacing NetApp, I would suggest paying for a full time engineer from the next storage company you buy from. They can manage the array for you (properly) and ensure these things don't happen. Otherwise you'll need your sysadmins to start reading product documentation, following best practices, and testing procedures.
The email is a careful balance of information that 90% of people will find useful and not too much information that no one understands it. Never once did we say we are not to blame, actually the opposite, it is our responsibility no matter the vendor or what we replace the hardware with.
Basically, this just boils down to a time issue. No data was lost it just took time for things be to rosy again. Restoring from backups or switching to a DR site takes time too. If you have never fully tested the DR site it might take a long time.
It it easy to be Cpt. Obvious; "Well, there's you problem right there" in cases like these but it just sounds like they need better documentation about what to do in the event something like this happens.
Forgetting about the DR site for a minute. Why is the NetApp a single point of failure?? Given they are extremely stable and running multiple heads further reduces this but if a single issue with one array causes massive downtime then you might want to think about a snap mirror to a second filer. Switching to a different storage vendor doesn't sound like it will fix the underlining issue here!!
I am very curious as to why a two disk failure caused an outage. What exactly happened when both disks failed?
The 2 disk failure did not cause the outage, but the process the filer head had to go through to get the data back onto new drives and then further actions taken with SnapMirror and other items to try and recover faster.
That being the case, the two disk failures didn't need to be concurrent for you to end up where you did...
I'd love to chat a bit about your experiences.. I'll even take you out to lunch (I'm right down the street in Newton). :P