I wish these sort of papers would focus more on the 3 percent that it got wrong. Is it wrong by saying that a day would have a slight drizzle but it was actually sunny all day, or was it wrong by missing a catastrophic rain storm that devastated a region?
I've worked on several projects trying to use AI to model various CFD calculations. We could trivially get 90+ percent accuracy on a bunch of metrics, the problem is that its almost always the really important and critical cases that end up being wrong in the AI model.