Show HN: Ranking weather models by how their forecasts turned out
nickleenders.github.io
A few things that surprised me: - AIFS performs very well, yet almost no commercial apps give you access to it or uses it in their blend (I suspect some do without disclosing it tho) - Foreca scores surprisingly well compared to other apps and raw models - ICON is very accurate around the mediterranean, but performs quite poor everywhere else
There's also a history page that scores each model back through its full archive (about 5 years for GFS) to see if forecasts have actually gotten better.
It's a static page and open-source: https://github.com/NickLeenders/verisky-scoreboard. Public models are scored in your browser against Open-Meteo's archive. Commercial scores come as small aggregates from my server, because those providers' terms don't allow redistributing raw forecasts.
It powers an app that does the same per location in more detail, link is on the page.
Happy to answer questions about the scoring method.