Here are some questions: - How many spare drives were in the system when the first and second drive failed? Netapp does not shut down volumes of storage if spares are in the system to take over for the failed drives. - How long was it really before those drives were replaced with a spare that could take over and rebuild? - Why don't you publish your DR plan and explain exactly where it didn't go as tested and planned? Point this out to show where the issue occurred that had been previously tested and shown to work properly.
I am not an employee of NetApp and not even a customer of NetApp. I have used it in the past and like all technologies it requires the care and feeding that is well documented in manuals they provide. And it requires good administrators to do their jobs to test and monitor things and ensure the resources are available for the system to work as designed.
Once you have heard both sides of the story the only thing you learn is that there are more the 2 sides to the story.
When these types of things happen the folks closest to them always leave out details or cannibalize the story so that they can be found blameless. I have been in the IT field for 15 years and seen it time and time again (with many technologies).
I read an email from Grasshopper this AM detailing to their customers what this issue was. It was so vague and left so much to interpretation that it really came across as whiny and misinformed. It was extremely unprofessional to apologize and then blame (without full explanation or root cause).
There's more to this than meets the eye. Believe me.
Since you all are replacing NetApp, I would suggest paying for a full time engineer from the next storage company you buy from. They can manage the array for you (properly) and ensure these things don't happen. Otherwise you'll need your sysadmins to start reading product documentation, following best practices, and testing procedures.