I’m not certain this is the case (who could be?) as, like the author, I’m swimming in the same sea and can’t detect the long, deep waves, just see the short ones I’m swimming though. But I hope I’m right, though I won’t live to see it.
I’m not certain this is the case (who could be?) as, like the author, I’m swimming in the same sea and can’t detect the long, deep waves, just see the short ones I’m swimming though. But I hope I’m right, though I won’t live to see it.
The experience during my PhD indicated to me many areas are ripe for big theoretical advances using the data collected over the past 50 years. The reasons why we rarely see those advances I think are along the lines of the article. I think very few people are even trying to make sense of the huge amount of published data. Let me give a specific example from my own experience.
Here's the preprint of a paper I wrote during my PhD: https://engrxiv.org/preprint/view/740 (The cert expired yesterday, unfortunately.)
In this paper, I basically compiled a ton of data from the open literature, showed that common beliefs about a certain map in my field were in serious error (some parts of the map were absent, most boundaries were basically flat-out wrong, etc.), and developed theory and regressions for various parts of the map. It became clear to me that this sort of work, while valuable, is strongly disincentivized.
Here are some reasons:
1. To the vast majority of tenured academics, the deep literature review I did looks unproductive. My PhD advisor repeatedly called me a librarian as-if what I was doing was not my responsibility. Note the emphasis on deep in the first sentence of this point. I went a lot deeper than most people writing review articles did, tracking down a lot of obscure documents in the process. And I learned a lot in the process too! In most fields I'm sure there are many great papers that were unjustly ignored.
2. Compiling data from the open literature tends to be viewed as not science among many. Seems to be a variant of "not invented here" syndrome: "not measured here". I get that a lot of the point of a PhD is to learn how to make certain measurements, sure, and I think that motivates a lot of the unnecessary measurements made. But using existing data should be an option. And I've never seen anyone explicitly say this, but it seems to me that many people in academia believe that when data is first published, all its value has been extracted.
3. The process of actually consolidating the data into a single database is time-consuming. It may be that in terms of number of publications over the course of a PhD, doing "incremental" research has higher ROI. Data compilation makes more sense over a longer time frame in terms of ROI. Note that I'm only saying incremental work has higher ROI in terms of number of publications. Data compilation has higher ROI in terms of actual scientific value in most every case as far as I'm concerned.
4. The reception I got for this work was fairly muted. The latest fad (machine learning applied to fluid mechanics in my case) gets way more attention, even though almost every time those models are quite brittle, particularly given point number 2: The only data used is new and rarely comparable in comprehensiveness to all the published data. Using more data would of course help the ML models generalize. I remember a conversation I had with someone about machine learning in physics. I told him that I don't need machine learning because I compiled a huge amount of data from the literature and have found that linear regression works great. His response was along the lines of "That's unfair!" No, it's not. That's how science should work. But it seems to me that many people simply target the latest fad without thinking too much about the lessons of the latest fad.
It seems to me that scientists can find value in analyzing previously published data, it just takes time to find the data, compile it, and analyze it. And given that someone would probably get more publications and attention doing something else, they tend to not do the right thing by analyzing previously published data.
I wrote some similar things in this previous comment of mine: https://news.ycombinator.com/item?id=31054992