Totally wrong. You cannot generalize such a statement because it depends on the micro and macro weather conditions. A very stable situation makes it very easy to forecast one week and beyond. On the other hand, there can be situations where you cannot accurately predict the next 12 hours (e.g. cold air pool).
And that is exactly where these AI models will break down. They will "shine" (or fool us) with how well they predict the stable situations, and will produce utter rubbish when the high, turbulent and dynamic weather fronts make prediction difficult.
But of course, if your weather is "stable" 80% of the time, you can use those shiny examples to sell your tool, and count on user forgetfulness to get away with the 20% of nonsense predicted the rest of the time.
You can do a backtest for any point in the past, as long as you only use the data that was available until 15 days before the day being predicted.
> The initial Google paper stated that the Google Flu Trends predictions were 97% accurate comparing with CDC data.[4] However subsequent reports asserted that Google Flu Trends' predictions have been very inaccurate, especially in two high-profile cases. Google Flu Trends failed to predict the 2009 spring pandemic[12] and over the interval 2011–2013 it consistently overestimated relative flu incidence,
One of the difficulties with using user data to understand society is that the company isn't a static entity. Engineers are always changing their algorithms for purposes that have nothing to do with the things you're trying to observe. For Google Flu Trends specifically here's a great paper
https://gking.harvard.edu/files/gking/files/0314policyforumf...
https://www.nature.com/articles/s41586-024-08252-9
See section "Baselines".
"The highly non-linear physics of weather means that small initial uncertainties and errors can rapidly grow into large uncertainties about the future. Making important decisions often requires knowing not just a single probable scenario but the range of possible scenarios and how likely they are to occur."
https://en.wikipedia.org/wiki/Weather_forecasting#Persistenc...
IIRC the Metoffice in the UK does/did pay bonuses to staff based on modelling exceeding that criteria in a calendar year.
Again, IIRC, in the UK the persistence forecast suggests something around 200-250 days of the year have the same weather as the previous day.
>> We use 2019 as our test period, and, following the protocol in ref. 2, we initialize ML models using ERA5 at 06 UTC and 18 UTC, as these benefit from only 3 h of look-ahead (with the exception of sea surface temperature, which in ERA5 is updated once per 24 h). This ensures ML models are not afforded an unfair advantage by initializing from states with longer look-ahead windows.
See Baselines section in the paper that explains the methodology in more depth. They basically feed the competing models with data from weather stations and predict the weather in a certain time period. Then they compare the prediction with the ground truth from that period.
Plot twist: they measure accuracy in predicting the weather 5 years in the past.