I think this is a case where the engineers can look at the numbers and calculate for a limited variety of well defined (and foreseen) scenarios, but that they continue to fail to have an intuitive sense of the power of the thing that they built. Given the complexity of the real world, they can "engineer" themselves into an undue sense of certainty about what will happen when they light a rocket of this power... and I think we see the result of that. As an aside, this is precisely where I think we'll some of the biggest failings for LLM based AI... it can get you some nice words to describe a situation and may be able to do all the calculations... but having that gut feeling that maybe there's more to a situation than the numbers would have you believe is still something that humans can do better (though not SpaceX this time around).
In closing, if I had a binary choice of "successful" or "unsuccessful" for this test... I'd still probably call it successful. But only just; it's a very qualified success given some of the test parameters SpaceX had set for themselves.
(of course we could find out about even more damage to "stage 0" which might shift my assessment... but we know that they're in the process of putting in a deluge system at least already, so I'm being a bit more charitable regarding expectations than I might otherwise be.)