During a conversation with an academic researcher (non-Computer Science) friend, when I brought up the topic of data sharing, especially in context of the infamous "replication crisis", they have their reasons not sharing. I'm loosely paraphrasing here, while trying hard not to misrepresent/misremember their exact views:
"I want to protect my data; I don't have enough time to present my data in a presentable form; and more importantly, they'll just steal my idea and go present it as theirs—and I might lose funding" ... and so on.
I can empathize with the academic pressure of "publish or perish". And not least of all, "need some food on the table, and roof over my head".
But I still wonder, there must be other effective ways to gently persuade a said researcher (especially in the 'soft sciences'—I'm not using the term derogatorily) on the importance of sharing data that allows reproducibility of a given experiment?
The replication crisis will continue until publishers incentivize replication.
What remains is the "small matter" of persuading the publishers to action.
And that's a good thing. A revolution in academic science led by Elsevier would be like King George leading the American Revolution. Real fundamental change isn't going to come from the profiteers who caused the problem in the first place.
If NIST required that you publish the data (say within N years to cover the concern about getting scooped on follow-up papers), and dinged you on future funding applications if you didn't meet their quality/reproducibility metrics, perhaps that would help to align incentives.
This is the same sort of idea as requiring research from public funding to be put in open-access journals so that the public can benefit from it.
But since I'm an academic layman, I can't punch any valid holes in the idea. :-)
Regardless, I hope this idea gets discussed more by the real stake holders. Perhaps you might to do a public write-up of the idea.
Finance has the exact same problem.
Of course, that ignores the analysis section, which is somewhat important.
But still, a vast improvement over today's way, where you may end up spending a lot of time only to get a vague reject.
But the 'variable under test' is definitely part of the hypothesis, and a study without a hypothesis is, not useless, but much less likely to produce a useful result.
There's a relevant xkcd:
If the (pre-registered) hypothesis was that green jelly beans cause acne, then this is at least an interesting result.
If you run this experiment with no particular hypothesis, and then decide on that basis that green jelly beans cause acne, this is just a setup for a later failure to replicate.
At best. At worst, no one bothers checking your results, the company stops selling green jelly beans due to bad publicity, and people who enjoy the green ones are deprived of them for no good reason.
In the proposed experiment, the hypothesis is, "Green jelly beans cause acne.", and the variable under test is "whether or not green jelly beans cause acne". The hypothesis is basically a statement of what variable is under test AND a prediction of what the results of the test will be.
What I'm saying is that the prediction part of the hypothesis is basically a statement of bias. It's useful in noting what an interesting result would be and what the biases of the researcher might be, but it arguably is counterproductive in that it takes the focus off of observing the variable with an open mind.
If we simply preregister the variable under test, that gives us everything we actually need from the hypothesis, and avoids the problem you describe: instead of testing 20 different variables and only publishing on the one that yields a low-P result, we test the single variable and that's it.
The reality is a lot of time spent by scientists is actually defining the problem properly (even after you think you’ve defined the problem properly), and preregistration studies would hinder the flexibility to adapt and change the direction given new understanding.
The idea of evaluating experiment design and committing to publish before seeing any data/results also strikes me as deeply flawed. Graduate students can sit around think up many deep questions and draw up tons of beautiful experimental designs. That is the easier part. The harder part is actually running the experiments properly and interpreting the findings correctly. If we are publishing experiment ideas, every graduate student would be submitting 100s of ideas to top tier journals. However, most of those experiments would end up producing inconclusive or utterly garbage results, which would provide little insight if published.
The other issue I see is that many interesting findings come out of serendipitous results that are not related to the initial experiment or hypothesis. You start out trying to answer one question and then stumble upon some else that is very interesting (so you then explore that). If you have already submitted one experimental design and are supposed to publish on that, what do you do if you stumble upon something else that is actually much more interesting? In my experience with research you don't often go directly from one question/hypothesis/method to an answer and then publish. It's much less linear than that with many dead ends along the way. It's just not something you can pre-publish.
All of that said, one area where I think we would agree is on publishing the results of well performed experiments with null results. I think it would be useful to the community if people published reports saying "we had this interesting question, did a proper experiment, and it turns out the hypothesis was null and there was nothing interesting there." In the current system, that will likely not get published. But, other researchers would benefit from seeing that result because they may have the same hypothesis and may waste time trying to explore the same dead end. A Journal of Null Results could be useful here. However, you do run into the issue of not knowing if an experiment failed because of a researcher mistake or error...
The problem we're trying to solve isn't poor experimental design, it's a failure to publish "boring" results--a bias against uninteresting results which should cause us to distrust any interesting result.
> The idea of evaluating experiment design and committing to publish before seeing any data/results also strikes me as deeply flawed. Graduate students can sit around think up many deep questions and draw up tons of beautiful experimental designs. That is the easier part. The harder part is actually running the experiments properly and interpreting the findings correctly. If we are publishing experiment ideas, every graduate student would be submitting 100s of ideas to top tier journals. However, most of those experiments would end up producing inconclusive or utterly garbage results, which would provide little insight if published.
An inconclusive result does provide insight. If you have 1 study showing vaccines cause autism and 37651 studies which are inconclusive, that starts to be pretty conclusive.
If the results produced are "utterly garbage", then that really shows us that the "beautiful experimental designs" weren't actually effective.
> The other issue I see is that many interesting findings come out of serendipitous results that are not related to the initial experiment or hypothesis. You start out trying to answer one question and then stumble upon some else that is very interesting (so you then explore that). If you have already submitted one experimental design and are supposed to publish on that, what do you do if you stumble upon something else that is actually much more interesting? In my experience with research you don't often go directly from one question/hypothesis/method to an answer and then publish. It's much less linear than that with many dead ends along the way. It's just not something you can pre-publish.
You publish the thing you agreed to publish, mention the interesting phenomena in the analysis, and apply to study it.
> All of that said, one area where I think we would agree is on publishing the results of well performed experiments with null results. I think it would be useful to the community if people published reports saying "we had this interesting question, did a proper experiment, and it turns out the hypothesis was null and there was nothing interesting there." In the current system, that will likely not get published. But, other researchers would benefit from seeing that result because they may have the same hypothesis and may waste time trying to explore the same dead end. A Journal of Null Results could be useful here.
A journal of Null Results already exists, and doesn't solve the problem because it's just reversing the publication bias.
> However, you do run into the issue of not knowing if an experiment failed because of a researcher mistake or error...
I'm pulling this out because it is worth addressing, as it fundamentally misunderstands the problem.
An experiment that gets a null result is not a failure. An experiment which fails to observe the variable under test is a failure. If science is working correctly, I would expect the majority of correctly-performed experiments to produce null results, because the obvious phenomena have already been discovered. The idea that "interesting result = success and null result = failure" is fundamentally unscientific and needs to be stricken from the human consciousness.
And, the way this bias plays out in practice is that you're sitting here wondering if an experiment produced a null result because of researcher mistake or error, but not applying the same skepticism to "interesting" results. This is the opposite of a rational position: as far as I know, there isn't a replication crisis for null results. It's the interesting results which are having a replication crisis. If you're suspicious that null results are the result of experimental error (as you should be) then you should be far more suspicious that interesting results are the result of experimental error.
> The problem we're trying to solve isn't poor experimental design, it's a failure to publish "boring" results--a bias against uninteresting results which should cause us to distrust any interesting result.
I do not see the value in a quest to do boring and uninteresting research and publish it, even if the experimental methods are exquisite. The goal of PhD research is to explore and find something new or interesting to add to the scientific body. Graduates already cannot read all the interesting papers in their fields. If you add 10x more boring papers, they are never going to be read and will just be a waste of journal editor and reviewer time.
> An inconclusive result does provide insight. If you have 1 study showing vaccines cause autism and 37651 studies which are inconclusive, that starts to be pretty conclusive.
You can't add up 37651 insignificant results and say they equal a significant result. If there was a study with significant results showing no link between autism and the vaccine, it would get published (many such studies have been published). If your study just shows nothing, then most won't care because they can't learn much of anything from it.
> If the results produced are "utterly garbage", then that really shows us that the "beautiful experimental designs" weren't actually effective.
That is one of my points. Many experimental designs are beautiful and seem great, but then when researcher do the experiment many different issues can come up yielding little to no useful data. If you had to evaluate (for pre publication) every graduate students experiment ideas, you would run out of time!
> You publish the thing you agreed to publish, mention the interesting phenomena in the analysis, and apply to study it.
No one wants to read about the uninteresting experiment. They want to read about the interesting discovery and want learn about that. So, you publish the study of the interesting part and don't bother publishing the boring experiment.
> A journal of Null Results already exists, and doesn't solve the problem because it's just reversing the publication bias.
I guess I believe a bias towards publishing interesting studies with significant results is not a bad thing. Publishing inconclusive, uninteresting, or trivial findings does not add much value, if any, to the scientific body in my opinion.
> An experiment that gets a null result is not a failure. An experiment which fails to observe the variable under test is a failure. If science is working correctly, I would expect the majority of correctly-performed experiments to produce null results, because the obvious phenomena have already been discovered. The idea that "interesting result = success and null result = failure" is fundamentally unscientific and needs to be stricken from the human consciousness.
I don't believe "null result = failure", I believe many null results are not interesting and likely are not useful if published because many experiments do not produce useful data or significant results for a variety of reasons.
That's part of the problem. Publications are seen as achievements. If you got accepted to a prestigious journal or conference, you can list this on your CV as an impressive "award"-like thing. A publication list is not just a list of "Look, this is the kind of stuff I've been working on, have a great read at it", but "Look, my research is so great it got accepted to all these fancy places!".
Publications are therefore unfortunately not merely about sharing new info with the research community but an award show. Ideally a publication would be the start of the conversation: "this is what we found, this is the method we propose, what do you think of it, community? Will you pick it up?" The test is then whether the ideas get adopted. But that's harder to measure. Citations try to approximate it, but it's a very crude approximation. A citation, as such, may mean tons of different things: e.g. a) a deep critique (Negative impact) b) being cited as part of a long block of "these other works exist, too" (Low impact) c) another work substantially based on the deep ideas of the original paper (High impact), d) being listed in a table for comparison, ala "we beat this other method", but no other discussion of the original paper (Low impact), etc.
If a publication was nothing more than a "hey, look, this is interesting", then I'd say, publishing mostly novel sexy results would be fine! After all, the surprising cases are those that teach us the most. However, as I said earlier, a paper is not only about "hey, this is interesting", but a "hey I want to advance my career", too. And in a twisted way of logic, I can agree that therefore we could put a band-aid over some of the problem by publishing (rewarding) systematic work with negative or boring results. But ultimately, this goes against the original purpose of papers, that is alerting the scientific community to potentially new information that we haven't known about before.
----
Ideally, to assess someone's scientific career, there would be at least one smart, attentive, impartial expert taking their time reading through the publications, taking notes, pondering, digesting it all, consulting other experts etc. However, this is too subjective.
Quantitative metrics seem superficially more objective and therefore egalitarian. The original idea is probably that if we just based everything on subjective judgement of scientific importance instead of publication count, there would be even more networking and friendship-based quid-pro-quo back scratching.
But everyone is overworked, and those who aren't, want to keep it that way. So nobody wants to put in the effort to actually interact with the deep content of research. It's too complicated and too opaque.
----
The problem is, flashy results are by their nature more attention-grabbing on all levels. It's not just some small perverse incentive. This is how all of us work, this is how history works, how everything works. The winner takes all, the rich get richer etc. We remember the Einsteins of history, those who just worked systematically and didn't find much aren't heroes. And if that's our bar, then people will do everything to look like they clear it. In any system, scientists would have to hype up their impact, it doesn't matter who makes the decisions.
Currently, universities want to employ researchers who will make a visible impact. Because that means attracting funding, but also attracting bright people from all around. Career-conscious researchers want to go to universities that help them market themselves well (good PR departments etc). PR is not only for laypeople as the audience, there is such a flood of research nowadays that even the experts of a small niche cannot keep up with everything happening.
----
The root of it is human nature, competition, deception, cliques, hierarchies. But the new about it is the scale of it, and the accompanying mechanization of it all. The idea that you can mass-manufacture innovation. That you can expect thousands upon thousands of researchers to make regular breakthroughs and, to say my field as an example, publish tens of thousands of novel AI-related ideas every year. It's related to credential inflation, and fake signaling: people with good academic track records got the good jobs and the respect, so people try to emulate that. Everyone tries to become the 1%, the rock star. And everyone wants to hire the 1%. So just like an evolutionary pressure, people try to appear like the successful. Soon enough the old signal doesn't work anymore. It used to be a high mark of educational level to have passed high school. Today that's a bare minimum. College used to be a meaningful differentiator. Now more than half of young people are "college educated" in developed countries. The next step is about becoming "researchers". Nowadays, having some publications is not a big differentiator. We see this also in title inflation like monkeying around in Excel is "data science" and "AI".
It's not just a monkey's paw. There is no central figure orchestrating it, asking the monkey's paw for more papers. It's a distributed system of agents acting in their self-interest. Nobody wants papers for papers' sake, they want to make defensible, justifiable decisions that will not get them fired and will pass satisfaction up opposite the path where the money is flowing, all the way to the CEOs, politicians and taxpayers.
I don't think that this is necessarily a problem. The problem is that these "awards" are awarded based on results rather than on work.
In my mind, a prestigious journal should be one where the studies have a high percentage of replication of their published work. A high-impact study which fails to replicate is just a popular lie, and has no place in science.
After all, the journal system was invented to solve distribution of papers. We have the internet now, so is there any need for the journal system any more?
Independent reviewers would/could easily step up to pick up the interesting papers and present a "feed" of the good stuff.
One angle of my fascination was how the [rating] baby was pretty much tossed out with the bath water. It was shouted down and denounced by journalists (as for example an adult filter which ironically ended up its only application) Some journalists described a perspective as if they had a kind of tenure. They were published for so long that the idea of a rating system was just offensive. We could argue that a good rating system would use existing talent for calibration but the real question to ask imho is: If PICS was so bad, what did we get in stead? Anon 5 star ratings? Thumbs up? HN points? Number of github saved games? To say it doesn't compete with publishing in journals is somewhat of an understatement.
End of the day all we are looking for is good meta data. If note worthy people in a field want to endorse a HN topic, a blog posting, a usenet post, a tweet, a youtube video, a facebook posting or a torrent[2] real credit could go to the author.
A rating system or spec therefore could simply accommodate that process. (It should for example require the author and their endorsers make backups available.)
Journals are from the horse and carriage days. It is quite embarrassing how we didn't come up with something modern.
[1] - https://www.w3.org/PICS/services-960303.html
[2] - torrents are nice to share huge data sets
I hadn't heard of PICS before, it's interesting. Is there a published story about why it failed anywhere?
It isn't just journalists apparently, I think most people hate anything that smells like scrutiny. There are probably plenty under appreciated authors. They ironically have no way of knowing. I'm convinced there are and always have been fantastic ideas out there that didn't even get written down. Why even bother having them or fleshing them out if they cant be judged?
In my inner dialog I consider it the greatest puzzle of all times (and expand the scope to rating all human creations) if anyone ever truly solves it all previous revolutions will be reduced to simple events. The potential for recursive self improvement of such system is probably equal to a general AI but the results will be better.
haha, with the lack of published stories I just discover I sound to myself like I have a lot of explaining to do making such fancy claims. I will have to ponder writing it into a blog. For now a business plan that everyone is specifically uninterested in is not going to work.
But then again, that is like, just everyone's opinion? We have no means of judging the value of it.
It'd probably make writing them easier too.
The problem is way deeper than just academia. Such as, is there fairness in the world deep down, is mass-produced excellence possible? Does individual greatness actually exist or is it all just a power play?
Overall, the quality of science is extremely difficult to measure, precisely because it operates on the border of the unknown and because people try their best to appear the best possible. Science is difficult to understand and is often far removed from the here and now, and may only bear fruit decades down the line. It's hard to judge for the same reason that antelopes are hard to catch for cheetahs: competition (mainly the antelope vs antelope type).
In the end, science has only been this mass product for a few decades. Before that it was mostly a pastime of weird nerdy aristocrats or people paid by aristocrats for showoff purposes. Or church people with too much time on their hands.
In reality, from the top down view it's a huge gamble. You try to get good people to do their honest best and then see what happens. Then at the end there will be some breakthroughs. But only a few every few years in each field. However, this does not satisfy the participants. I toiled away as well, but the reward is only paid to the lucky one. So everyone tries to be the lucky one, which perversely pushes everyone to take fewer risks, making the collective likelihood of a breakthrough lower, but their own expected reward better.
-----
My grandfather used to recite the story of a farmer who had three pigs. Every morning he'd throw two apples in their confinement. He'd then grab a big stick and beat the one that didn't get any apple: why didn't it try harder?
-----
My prediction is that as with all signaling spirals and treadmill effects, there will be something new to aspire to, to tell the wheat from the chaff, a signal that's harder to fake. It's a constant race. You demonstrate your fitness by being adaptive to how the system changes. Overall the "quality" of people obviously doesn't change over time, it's just that the competent/powerful drive the criteria to their benefit.
As academia/publishing etc. is now flooded with "the plebs", "the elite" will move on and will perhaps use other criteria.
----
Now, going back to assuming this is about the object-level science itself. Where to find the best science? You cannot do this in general. You have to educate yourself and dive in yourself. You try to learn how to judge people's character and try to listen to and digest the assessment of those you trust.
There's no other way, gather experience and become "better" yourself. Use the cognitive resources of your brain to try and outsmart your opponent: the writer of the piece of text you are reading. This cannot be standardized/metrificated in a simple way (outside human-level AGI). If your organization does not put in the cognitive power of extensively processing the content of a particular research and critically examining the motivations behind it etc., there is no way to judge it. Then you're back to credentials. Did it come from a highly cited person? Is this person endorsed by other big shots, where the "seed big shots" are the researchers at the historically most prestigious institutions.
----
Currently, to find interesting research I personally use Github recommendations, Google Scholar alerts watching for citations of landmark papers (good indicators for progress in niches) and authors. A well-curated Twitter-feed is also useful, as is arxiv-sanity. In the end, I have to make up my mind if it's good work or not, and as everyone I don't have infinite cognitive resources. So I make snap judgements based on paper gestalt, affiliations, plausibility, result tables, etc. If it clears this bar, I dive in more. Over time, you learn to trust some smart people and can follow them online and see what they say and recommend. And continuously learn and grind your brain. Cognitive work cannot be spared, just like you cannot spare physical exhaustion in sport competitions.
Well, I think there's two separate problems:
1. Academic journals not publishing "boring" results, which incentivizes scientists to bias toward novelty and not publish null results. I think the solution to this is for journals to accept based on methodology and subject matter of experiments/studies, and to commit to publish before the study is even done, so there's no chance that results can influence the decision to publish.
2. Science journalism being written by people unqualified to analyze the science who instead sensationalize for clicks.
I'm not particularly concerned that the people are "only reading papers from big shots". The problem is that prestige is measured by impact analysis rather than replication. If you incentivize a metric which is largely a popularity contest for ideas, it should be no surprise that what you get is popular lies[1]. One might think that the focus on impact in prestigious journals means there might be higher replication in less prestigious journals, but from what I can tell, the less prestigious journals are no better: they're mostly targeting the same metrics, they're just less successful at reaching their goals.
[1] My intent is not to actually accusing anyone of intent to deceive here--I just can't quite figure out the right word. Untruths doesn't quite match what I'm trying to say.