Deep learning for twelve hour precipitation forecasts
nature.com
nature.com
This is the case with this paper, from which I quote (see section "Dataset creation and splits"):
>> The available data spans a period from July 2017 to August 2020.
In other words, they have shown that their system works better than an earlier system on old data. It can predict rainfall in a three-ish year interval in the past.
Show that your approach works on real-world data, that you didn't have access to when training your model. Show that your models can predict the future, not the past. Otherwise, your experiments don't mean, what you think they mean.
Edit: and to make it perfectly clear for the uninitiated: when you, as an experimenter, know the "ground truth" of your test data, there is nothing easier than to tune the training of your model to maximise its performance on the test data. That's doubly so, and doubly as dangerous, with neural nets that are very, very good at reproducing their training set, but very, very bad at generalising beyond it. Unfortunately, this is what the vast majority of experiments in machine learning do: they test on known data. The result is that nobody really knows how good the tested systems are until they deploy them in a real environment (if they ever do, which they usually don't, because the whole point was writing a paper to report improved performance, and then move on to the next).
The idea is that you don't even look at this data until you're confident you've solved the problem, and then you have one shot to confirm that you have, indeed, solved it. Obviously you have to be careful in selecting this set -- if you're working with time-series, for example, you want it to be temporally disjoint -- but done well it's a reasonable approximation for "future results".
About as good as you can get without collecting new data.
I can't tell if they did that here, but if they did, they didn't use the common jargon for it.
Human nature, and publish-or-perish rewards, mean the researcher will say "well, what if I try this other model? Or tweak this parameter, or ..."
In other words, it's unrealistic to say you have one shot to confirm that have, indeed, solved it.
Do you really believe that this practice is unrealistic because anyone with a model that doesn't do well on the held-out set will simply engage in fraud? That is a very, very cynical take.
FWIW, I have published this way [0]. There were definitely some aspects of the model's performance on the held-out set that did not meet test set performance; it's not fatal to the paper, it's an invitation to discuss.
Fraud is not needed. My intuition is that most people publishing in machine learning have simply not thought carefully enough about their experimental methods. They're not defrauding anyone, they're misunderstanding their own results.
And since the majority of published papers do the same, it's very difficult to convince anyone what the good practice is.
Good on you for doing it the right way.
But explicitly holding out a test set, throwing it away because your results on that set were no good, trying again, and then representing that held-out set as, indeed, held-out in later work -- as the parent suggested? I'm not sure what other word to use there.
A common split is train/validate/test, but all three are used during training -- train to actually train, validate for intermediate loss, test for model comparison.
What you want is a fourth, held-out test set that isn't looked at until publish time.
This paper has two test sets, but they have different data properties, and it's not clear they were held out until publication.
So let’s call holdout datasets “second-best practice”?
Sadly too common that collecting new data is practically speaking impossible, because the actual collection process is underspecified.
"The training, validation and test data sets are generated without overlap from periods in sequence. Successive periods of 400, 12, 40, 40 and 12 h are used to sample, respectively, training, validation, and test data, with the two 12 h periods inserted as hiatus."
I know nothing about machine learning, but I was in grad school with others studying machine learning and I took a Coursera course on the subject, and carving out a training set is utterly standard practice. A paper that didn't do this would get filtered by a grad student reviewer. So I'm not sure what possible sub-category of machine learning your statements could apply to. Can you share any examples? For example, do any of the famous papers in the field fail to hold out test data? Something like AlphaGo or AlphaZero is indeed testing on all new data -- new games it plays with others. Do any papers in well-known machine learning fora fail to hold out test data? Does anything that has bubbled to the front page of HN? It would be interesting to look at if so.
Do you mean p-hacking? That's a much more subtle thing.
When you use your test set as part of model selection, you don't really have an empirical test set -- you do have a training set that you've partitioned into various parts for various purposes, but you don't have a proxy for how the system will perform in the face of new data, because you're reporting results that are based on data that you used as part of your model-generating process.
If you try out models A, B, C, and D while developing, and then it turns out model D has the best performance on your test set, and you report your results using model D and that same test set, you are essentially "training" using your test set.
Yes, in some problem contexts, for ML researchers who are also in the business of generating their own datasets, the data they test on can be entirely new. But for anyone using a pre-existing dataset, holding out a test set for your project is not yet a well-established practice -- and this makes it somewhat harder to trust the results.
Edit: Oh, and I see now that you explained it below, sorry.
Great, but ... that requires you have collection infrastructure and almost by definition means you can't make it work on a dataset. The way ML research works is like the Kaggle competitions.
You get a dataset. This can be anything, as long as it can be expressed as a tensor (which are vectors, but can have more dimensions). Then split into seen and unseen data. You split into 3 portions/partitions. One for training. One for testing, and one for blind validation (80% training and 10% each for test and validation sets).
In order to simulate what you're suggesting researchers can, and often do, take the most recent datapoints as validation and testing sets.
But extending the data set is an expensive process that generally can't be done by the machine learning researchers themselves. So it is not done.
:)
> A family father gets premonitions about a massive storm that will come, which triggers him building a bunker in his back yard. It is a family drama on how his obsession with protecting his family slowly tears it apart. In the end they convince the father to go on vacation for a few days, just to get away from the obsession. As they arrive in their vacation home, a huge, black storm cloud rolls in on the horizon
Great Movie, recommend watching it.
Anyway, everyone just ignored it. Except the scientist, who I suspect FTF out.
New series of propaganda shows to get people to empathize with AI.
It doesn't cause any such discomfort, because in the grand scheme of things, while this is still very impressive and interesting work, it's still a toy in comparison with what tool like the HRRR is typically used for. Reported increases in performance are skewed towards the first few hours of the forecast, where all CAMs generally have some issues because of inconsistencies between the assimilated model initial state and the real-world (e.g. small deficits in the structure of convective systems at the initial state can dominate precipitation forecast skill at short lead times), and where traditional nowcasting systems are already significantly superior.
There's little evidence reported that this modeling paradigm can even fundamentally tackle the most critical aspects of short-term mesoscale/convective forecasting, which is hysteresis - the initiation of convection and the structure it takes on in different environments. This is _by far_ the most important way that the HRRR is used to aid in short-range forecasting.
In the long arc of things, the community is very, very excited to see how novel approaches involving things like generative AI could lead to next-generation warn-on-forecast systems and large ensembles. And while this paper is a cool early step in this direction, it's a very small one in the bigger picture.
I'm imagining those weather stations you can buy for your house coming with a GPU and a local model you can finetune on your own measurements and the measurements of people nearby - get an ensemble of predictions from neighboring models... Seems like a neat addition to a smart home.
Trying to run your own model with nearby inputs is kind of pointless because to actually assimilate measurements in your neighborhood, you'd have to run at an absurdly high resolution that makes it way too expensive to run an operational forecast as a commodity. You'd be limited in the size of the forecast domain, so within a few hours (maybe a day at most), your forecast would be dominated by boundary conditions from the parent regional/global models you force it with, and there would be no further propagation of information from your local observations to refine the forecast.
Deep learning can reduce the amount of compute for the same quality of results.
Deep learning applications in the field haven't come anywhere even _close_ to tackling these niches of the field yet. SOTA DL-based forecasting tools run at quarter-degree resolution, if even that, and the field hasn't even begun to run hierarchical or multi-scale models coupling coarse models to mesoscale or finer-grained. Hell - you'd be hard-pressed to even find mesoscale DL weather simulations in the first place! So it's a big, big stretch to suggest that DL can achieve "the same quality results" when no one has even offered a cursory glance at an AI application bordering on the SOTA in NWP.
There's an analogy to the 90s dream of replacing fluid dynamics with its differential equations with cellular automata. The claim made about the former is that the "variables in the equations represent real things in the world", but in the end it is all just a story. If the cellular automata/neural modelling can tell a better story ....
This canard always gets thrown out but the application here is flawed in two very big ways. First, "chaotic behavior" does not mean "unpredictable behavior." Modern numerical weather forecasting already deals with this through ensemble modeling techniques and other approaches designed to capture the statistics of the evolution of the weather, not just a single deterministic state. Anyone selling you a deterministic weather forecast from a single model is robbing you.
Second, DL will suffer the same challenges here because many AI-based weather forecasting tools are auto-regressive, where a model output is used to seed the next step of the forecast. So the AI approach doesn't actually escape this hypothetical limitation (in fact it might compound it badly).
Uhm no, it is all fluid dynamics. You are correct that weather is highly sensitive to initial conditions, but you are incorrect to conclude that that means DNNs are somehow better at dealing with that.
The only help they might be is in correcting systematic modelling and measurement errors, or speeding up computation (by guessing).
"The universe is made of stories, not atoms"
FD is just one of the stories we have to tell that allows us to predict and therefore engineer certain aspects of the universe in which we find ourselves. But it's no more than a story, no matter how fit it may be for the purpose.
Also, the claim would not be that DNNs are better at handling initial conditions, but rather than they are better at spotting patterns (ok, ok, call them correlations) than either ourselves or the systems we've built based on pre-conceived physical models.