A new replication crisis: Research that is less likely to be true is cited more
ucsdnews.ucsd.edu
ucsdnews.ucsd.edu
My whole perception of academia and peer review changed that day.
Edit to elaborate: like many of our institutions, peer review is an effective system in many ways but was designed assuming good faith. Reviewers accept the author’s results on faith and largely just check to make sure you didn’t forget any obvious angles to cover and that the import of the work is worth flagging for the whole community to read. Since there’s no actual verification of results, it’s vulnerable to attack by dishonesty.
In most cases, peer reviewers will just assume that authors claiming the "code is available" means that a) it is reproducible and b) it is actually there.
As a counter example, this recent splashy paper
https://www.nature.com/articles/s41587-021-00907-6
claims the code is available on github, but the github version ( https://github.com/jameswweis/delphi ) contains the actual model only as a Pickle file, and contains no data or featurization.
So clearly, the peer reviewers didn't look at it.
I think it's more important for reviewers to read the source, the same way one would read an experimental protocol and supplementary information, mainly checking for discrepancies between what the paper claims is happening and what is actually being done. In the above example, a reviewer reading the code would have spotted that the model isn't there at all, even though it runs fine.
In fact, I'd actually go further and question what kinds of errors could possibly be caught be running the same software that the authors did? Any accidental bugs will remain, and any malicious tampering with the experiment data is exceedingly unlikely to be caught even with a careful audit of the code.
One perspective is that, “knowledge generation wise,” the current system really does work from a long term perspective. Evolutionary pressure keeps the good work alive while bad work dies. Like that [Top Institution] paper: if nobody else could reproduce it, then the ideas within it die because nobody can extend the work.
But that comes at the heavy short term cost of good researchers getting duped into wasting time and bad researchers seeing incentives in lying. Which will make academia less attractive to the kind of people that ought to be there, dragging down the whole community.
Data manipulation generally doesn't happen by changing values in a data frame. It's done by running and rerunning similar models with slightly different specifications to get a P value under .05, or by applying various "manipulations" to variables or the models themselves for the same effect. It's much easier to identify this when you have the code that was used to recreate whatever was eventually published.
Don’t (credible) journalists have an honour system of getting at least three sources for a story?
Can’t we make researchers get at least two more confirmations from separate teams for something far more important?
If publication would require two more confirmations from separate teams, that would mean (a) doing the work in triplicate, so you get three times less results for the same effort; (b) the process would take twice as long as I spend a year doing the experiment and then someone else can start and spend a year doing the same experiment, and only then it gets published; (c) there's a funding issue - I have somehow got funding to spend many months of multiple people on this, but who's paying the other independent teams to do that?; (d) it's not a given that there are two other teams capable of doing the exact same research, e.g. if you want to publish a study on the results of an innovative surgery procedure, it's plausible that there aren't (yet!) any other surgeons worldwide who are ready to perform that operation, that will come some time after the publication; (e) many types of science really can't get a separate confirmation - for example, we have only one Large Hadron Collider, you can't re-do archeological digs, event-specific on-site sociological data gathering can't really be repeated, etc; so you have to take the data at face value.
Hopefully it is clear that that data is useless without some written text explaining what it means. Given that for hundreds of years the accepted way of presenting that explanatory text was by writing papers, I don't see any reason to abandon that. Tweaking our strategies for replication (after a description of the experiment has been published!) and reputation don't seem to contradict that.
For me a better solution would be to properly incentivise replication work and solid scientific principles. If repeating an experiment and getting a contradictory result carried the same kudos as running the original experiment then I think we'd be in a healthier place. Similarly if doing the 'scientific grind work' of working out mistakes in experimental practice that can affect results and, ultimately, our understanding of the universe around us.
I think an analogy with software development works pretty well: often the incentives point towards adding new features above all else. Rarely is sitting down and grinding through the litany of small bugs prioritised, but as any dev will tell you doing that grind work is as important otherwise you'll run in to a wall of technical debt and the whole thing will come tumbling down.
You have big companies making Billions with the work of relatively poorly paid nerds. But as soon as you make it possible for the nerds to claim all the profits of the work then you have a whole class of people whose job is to insert themselves as middlemen and ruin it for everyone, both customers and developers.
So basically the aim is to limit the degree to which you can privately profit from science, and expand the amount of science you can easily build on. You still get enough incentives for progress, the benefits accrue to society as a whole, and competition and change is enabled without powerful gatekeepers controlling too much in their own interests.
In the past people who did science could do so with less personally on the line. In the early days you had men of letters like Cavendish who didn't really need to care if you liked what he wrote, he'd be fine without any grants. That obviously doesn't work for everyone, but then the tenure system developed for a similar reason: you have to be able to follow an unproductive path sometimes without starving. And that can mean unproductive in that you don't find anything or in that your peers don't rate your work. There'd be a gap between being a young researcher and tenured, sure.
Nowadays there's an army of precariously employed phds and postdocs. Publish or perish is a trope. People get really quite old while still being juniors in some sense, and during that time everyone is thinking "I have to not jeopardise my career".
When you have a system where all the agents are under huge pressure, they adapt in certain ways: take safer bets, write more papers from each experiment, cooperate with others for mutual gain, congregate around previous winners, generally more risk reducing behaviour.
Perhaps the thing to do is make a hard barrier: everyone who wants to be a researcher needs to get tenure after undergrad, or not at all. (Or after masters or whatever, I wouldn't know.) Those people then get a grant for life. It will be hard to get one of these, but it will be clear if you have to give up. Lab assistants and other untenured staff know what they are negotiating for. Tenured young people can start a family and not have the rug pulled out when they write something interesting.
A better solution would be to stop overproducing PhDs. We could reduce funding for PhD students and re-direct that towards more postdoctoral positions - perhaps even make research scientist a viable career choice?
The act of producing a doctoral dissertation usually leaves something of a mark on one's outlook, skills, etc. I claim it is a _distinguishable_ achievement for life.
I'm suggesting that we re-direct some of the funding for training PhD students into funding for postdoctoral positions (via either fellowships or research grants). Professors would still get their research team, but rather than consisting mostly of untrained PhD students, they'd have a smaller, but more effective team of trained researchers.
Immediately after undergrad is how it used to work in the golden days of science, more or less.
If the competitiveness is the problem maybe tenure should be a lottery that you enter once at a fixed stage, preferably before you're expected to start publishing in journals.
A tenure lottery seems like an extreme option - there has to be a middle ground between what we have now and something entirely random.
I personally favor requirements which call for bundling raw datasets with the "papers". The data storage and transmission is very cheap now so there isn't a need to restrict ourselves to just texts. We should still be able to check all of the thrown out "outliers" from the datasets. An aim should be to make the tricks for massaging data nonviable. Even if you found your first data set was full of embarassing screw ups due to doing it hungover and mixing up step order it could be helpful to get a collection of "known errors" to analyze. Optimistically it could also uncover phenomenon scientests thought was them screwing up like say cosmic background radiation being taken as just noise and not really there.
Paper reviewing is already a problem but adding some transparency should help.
You don't have to convict people for full-on fraud. If you are caught using an obvious mistake in your favor or using a weak statistical approach, the punishment can be you are not allowed to apply for grants with a supervisor/co-PI/etc who's role is to prevent you from following that "dumb" process in the future.
Something like a well funded ten year campaign to do peer review, retrying experiments and publishing papers on why results are wrong.
I have a co-worker who had a job than involved publishing research papers. Based on his horror stories it seems like the most effective course of action is to attack the credibility of those who fudges results.
There will always be cases of fraud if someone deeps deeply enough into large institutions. That doesn't actually indicate that there is a problem.
Launching in to change complex systems like the research community based on a couple of anecdotes and just-so stories is a great way not actually achieving anything meaningful. There needs to be a very thorough, emotionally and technically correct enumeration of what the actual problem(s) are.
Research is heavily funded because people believe it's something more than a random claim making machine. You say governments should assume research is wrong and then try to replicate any claim before acting on it. But you end up in a catch 22: if the research community is constantly producing wrong claims there's no reason to believe your replication attempt is correct, as it will presumably be done by researchers or people who are closely aligned.
Additionally inability to replicate is only one of many possible problems with a paper. Many badly designed studies that cannot tell you anything will easily replicate. A lot of papers are of the form "Wet pavements cause umbrella usage". That'll replicate every single time, but it's not telling you anything useful about the world. Merely trying to fix things with lots of replication studies thus won't really solve the problem.
"Wet pavements cause umbrella usage" is something where I'd want to see your specific examples because it's easy to get a correlational study of that nature but very hard to design a causal one. The correlational studies are usually accurate and often useful for other research.
So a lot of people only notice this in the rare cases when someone within the academy decides to write about it. This can make it seem like science is self correcting, but it appears in reality it's not. When measured quantitatively there is no real improvement over time. Alvaro de Menard has written extensively on this topic and presented data on the evolution of P values over the last decade:
https://fantasticanachronism.com/2020/09/11/whats-wrong-with...
Additionally as he observes at the end of his essay, the problems are due to bad incentives, so the only true changes can come from changes to incentives. However those incentives are set by the government. Individual scientists cannot themselves change the incentives. The granting agencies are entirely oblivious to the problems and the scale of their ambition is in no way equal to the scale of their problem:
"If you look at the NSF's 2019 Performance Highlights, you'll find items such as "Foster a culture of inclusion through change management efforts" (Status: "Achieved") and "Inform applicants whether their proposals have been declined or recommended for funding in a timely manner" (Status: "Not Achieved") .... We're talking about an organization with an 8 billion dollar budget that is responsible for a huge part of social science funding, and they can't manage to inform people that their grant was declined! These are the people we must depend on to fix everything."
Firstly, the problem here is not an epidemic of scientists who feel too financially insecure to do good work. Many of the worst papers are being written by people with decades-long careers and who lead large labs. Their funding is very secure. They are doing bad work anyway for other reasons, sometimes political or ideological, more often because doing bad work results in attention, praise and power. Or sometimes because they don't know how to explain their chosen question, but don't want to admit that scientifically they failed and don't know where to go next.
Secondly, as you already realized your proposal relies on identifying which scientists have a proven track record, but the whole problem is that science is flooded with fraudulent/garbage claims which are highly cited ("proven") and which were written by large teams of supposedly respectable scientists at supposedly respectable institutions. Any metric you can invent to decide who or what has a proven track record is going to be circular in this regard. To Rumsfeld the problem, we are surrounded by "unknown knowns". You say this is an open question but to me that's a fatal flaw.
So the problem is actually the inverse. You say at the end, well, scientists who can fund their own work are an exception. Obviously in most cases scientists don't need to do this, they can also be funded by companies. Most computer science research works this way. Better CPUs and hardware is done almost entirely by companies. AI research has been driven by corporate scientists, and so on. In contrast academic funding comes primarily from government agencies that distribute money according to the desires of academics. This means a tiny number of people control large sums of money, and they are accountable to nobody except themselves. There are no systems or controls on academic behavior except peer review, which is largely useless because the peers are doing the same bad things as everyone else.
Viewed from an economic perspective academia is a planned reputation economy. The state is the source of all resource allocation decisions (academics being effectively state employees in most fields). There's also a deeply embedded Marxist worldview: universities have no working mechanisms to detect fraud, because of an implicit assumption that deep down when market forces are gone everyone is automatically honest and good. The hierarchy is stagnant; the same institutions remain at the top for centuries. A good reputation lets them select the people with the reputation for being smart (e.g. by school grade), so that reputation accrues to the institutions, which lets them keep selecting intake by reputation and so on. Supposedly Oxford and Cambridge are the best UK universities, they always have been, and they always will be. In a competitive, free market economy they would face competition and other institutions would seek to figure out what their secret is and copy it, like how so many companies try to copy the Toyota Way. In science this doesn't happen because there's nothing to copy: these institutions aren't actually different.
This implies a simple solution, just privatize it all. It would be wrenching, just like it was when the USSR transitioned to a market economy, just like it was when China (sort of) did the same. But one thing the 20th century teaches us is that you can't really fix the problems of a planned economy by tinkering with small reforms at the edges. The Soviets weren't able to fix their culture with glasnost and perestroika. They eventually had to give up on the whole thing. Replacing the current reputation economy with a real economy, with all the mechanisms that economic system has evolved (markets, prices, regulators, court cases, fraud laws etc), seems like a more direct and obvious approach to making things better, even if it may sound extreme.
My envisioned solution is similar to yours, here. But rather than "privatize science", which I think most people will interpret as "move to industrial research", my rallying cry is a little more like "hey scientists, stop depending on public funding, let's find creative ways to get the science done."
I also like to point out that money is often not the missing factor as much as community. This has always been true. Mendel discovered genetics by experimenting on beanstalks in his garden at his monastery. It cost him very little to do it, and he only stopped the research when his community told him to stop wasting time on beans and get back to the important accounting work that impacted the church's politics at the time.
You might think that maybe science was cheap in the past, but that today you need lots of money, to get the lab equipment, etc. However, science always has a cutting edge of cheaply evaluable questions. We recently hosted a DIY Synthetic Biologist (currently on the homepage of https://invisible.college) who showed the actual costs of his work, and his laboratory equipment was far, far, cheaper than the "cost" of his time. We can get far more science done with "amateur scientists" (remember that "ama" means love, and an amateur scientist is one doing science for love) by creating a scientific community outside the institutions for interested parties to work together, pool their brainpower and resources, and come up with great novel work.
And if anyone else agrees with me on this, please let me know so we can forces. I'm toomim@gmail.com, and am doing work on invisible.college.
I absolutely agree that a lot of science can be done very cheaply. Some of the most impactful papers were done by people who weren't in an institutional framework, even in the modern era (Satoshi being an obvious example). Additionally it seems most of the really problematic fields are ones where the budget gets dispersed over large number of people writing very cheap low budget papers, hence millions of social science papers with tiny sample sizes.
I'm a big supporter of industrial research though. Many great papers come out of industrial labs. Modern computing is practically defined by such research. The big advances all seem to come from big corporate labs (Xerox PARC, Bell Labs, Google, DeepMind, IBM, Sun, Microsoft, etc). The research is powerful because it's funded by people who expect some sort of meaningful results and supervise the work to ensure it doesn't go completely off the rails. Academic institutions have developed this totally hands off attitude that makes research more or less unaccountable to any standard beyond "will it get published", which in turn can be rephrased as "are the claims interesting".
> The big advances all seem to come from big corporate labs
That's an interesting claim, and I'd encourage you to find some statistics to verify this hypothesis, because in my experience, that doesn't ring true.
From my subjective perspective, it seems that academic and industrial research labs innovate at roughly the same rate per-capita. I was a PhD student when Microsoft was dominant, hiring the best faculty from all top-4 CS schools (CMU, Berkeley, MIT, Stanford), and they certainly produced a lot of papers, and did seem to dominate conferences, but the actual innovation in computing came from Apple and startups, which did not have "research labs". Microsoft, including its giant industrial research lab, certainly was not the driver of innovation in computing!
And here are some numbers to back that up: Microsoft's R&D budget in 2011 was 10x the budget of the entire NSF -- for all sciences. Yet, Microsoft was clearly not producing more than 10x the scientific output of all NSF-funded academic science.
So it would help to have some statistics for the claim that industrial research innovates more than academic research. They certainly pay more, and often hire more people, but per-capita they don't seem any more productive or healthier than academics.
Apple does very little research, in the conventional scientific sense we're discussing here, I think that's pretty uncontroversial. They produce few if any papers. They are (or were, under Jobs) very good at coming up with new ideas that strongly appeal to the buyer and which got them a reputation for innovation, but which probably wouldn't be considered clever enough to be research papers. At least not top tier papers.
For example, exposé is a widely imitated feature and was considered very innovative at the time, but it wouldn't be seen as serious computer science. The iPhone is/was widely considered innovative but had basically no new research tech in it, given that capacitive touch screens weren't developed by Apple. It was just a really nicely implemented mobile computer. Actually the innovations in the iPhone are nearly all packagings of tech developed by third party firms that Apple then buys or buys exclusivity rights too. At least, that's true in my view.
Microsoft's R&D budget I think is also a victim of definitions. Software firms normally report all product development as R&D, right? I think these days they may even report datacenter builds as R&D. We can see this on Microsoft's investor website:
"In addition to our main research and development operations, we also operate Microsoft Research. Microsoft Research is one of the world's largest computer science research organizations"
i.e. the kind of university type "scientific" research we're discussing here is only a sideshow in Microsoft's R&D budget.
You're right to call me out though; I don't have any stats to prove that industrial research does more than academic research. It's not a statistical argument to begin with, just my own own perception ("all seem to"). I read a lot of CS papers and the best ones have corporate email addresses at the top - the second best, a mix of corporate and university addresses, the third best, only university addresses. If you asked the man on the street to name the biggest innovations in computing in the past 20 years they'd probably say things like, uh, smartphones, YouTube, AI, blockchain, etc etc. All things that have little connection to universities, with AI being the closest but it was Google that revived that whole field and has been pushing it forward ever since. Neural nets weren't receiving much investment by the academic community before that.
Anyway, that's CS. CS really isn't the problem here. The pseudo-science is elsewhere.
https://nintil.com/newton-hypothesis
https://news.ycombinator.com/item?id=25787745
Due to career and other reasons, there is a publish or perish crisis today.
Maybe we can do better by accepting not everyone can publish ground breaking results, and it's okay.
There are lots of incompetent people in academia, who later go to upper positions and decide your promotions by citation counts and how much papers you published. I have no realistic ideas how to counter this.
When I was in graduate school papers from one lab at Harvard were know to be “best case scenario”. Other labs had a rock solid reputation - if they said you could do X with their procedure, you could bet on it.
So basically we treated every claim as potential BS unless it came from a reputable lab or we or others had replicated it.
We need to create new a social institution of Anti-Science, which would work on other stimuli correlated with the amount of refuted articles. No tenures, no long-term contracts. If anti-scientist wished to have income it would need to refute science articles.
Create a platform allowing to hold a scientific debate between scientists and anti-scientists, for a scientist had an ability to defend his/her research.
No need to do anything special to prosecute, because Science is a very competitive, and availability of refutations would be used inevitable to stop career progressions of authors of refuted articles.
https://www.nature.com/news/failed-replications-put-stap-ste...
When I got into university and started alternating studying and work, I realised just how incredibly clueless even adults are. The "let's just try something and hope nothing bad happens" attitude permeates everything.
It's really a miracle the civilisation works as well as it does.
The upshot is that if something seems stupid, it probably is and can be improved.
Robocall scams are very high on the profit:human misery scale, but their hardly going to end civilization. Pollution, corruption, theft etc all make things worse, but we never see the better world without such things so it all feels very abstract. Of course you need to lock your doors etc that’s just the way things are.
So after a certain time spent, you are left with a choice of ‘massaging’ the data to get some results, or not and getting left behind those that do or were luckier in their research.
Ultimately, this is the problem.
For example, I imagine that archeological work is extremely high impact if excavation efforts led to discovery of ancient city.
Archeology paper would probably be less interesting if the paper said “we dug this area, found nothing”.
If one were to judge those two papers, obviously the discovery paper is higher impact than the negative result.
Not as valuable as a discovery, but very far off zero value. Yet the reward in academia would be near-zero.
That can end up being just as time consuming as doing the research to begin with. Often there is no time and no money to go back and do that. If your 'budget' is 6 month you're going to spend 6 month trying to get your experiment to work. You're not going to 'give up' after 4 month and spend 2 month putting together a "why we failed" paper.
Very valuable lesson, although it sure did suck at the time.
He got criticized for it.
A lab exercise like that could really just be selecting for chutzpah (feeling charitable) or arrogance (less charitable).
It has to work well enough to… work… and reproduce. That’s it. It’s not “survival of the fittest.” It’s “survival of a randomized subset of the fit.”
There’s even a set of thermodynamic arguments to the effect that systems are unlikely to exceed such minimum requirements for a given threshold. For example, if we are visited by interstellar travelers they are likely to be the absolute dumbest and most dysfunctional possible examples of beings capable of interstellar travel since anything more is a less likely thermodynamic state.
So much for Star Trek toga wearing utopian aliens.
Otoh, they would be aware about that, and they might have spent some time improving how genes (or what they have) and evolutionary selection works for them, so that, say, their species with time becomes brighter and brighter than what's actually needed. If they wanted to do that.
Also more intelligence does not equal better ideas. The world is full of crazy or amoral people with apparently very high IQs. Your average flat Earther probably has an above average IQ.
Improvement is a war against entropy and n^n^n^… combinatorics any way you slice it.
Slowly across hundreds and thousands of generations.
By adding evolutionary pressure, for what they want -- it'd be up to those space traveling aliens to decide -- they can change their species, generations into the future.
> Improvement is a war against entropy ...
Reasoning in that way, the humans would not have gotten brighter than the chimpanzee monkeys. There's been evolutionary pressure for the humans to get brighter, and it would be possible for you (I mean the humans), or the space travelers, to add artificial ev. pressure.
Anyway never mind all this, maybe talking about space travelers and the humans and their genes isn't the best way to spend the day. Have a nice day btw
I think this all the time.
One problem is PhD degrees are too costly to those who don't get academic or industrial success from them. But as long as talented people are willing to try to become a professor I don't see the system changing.
I think many more are drawn to professorship for a sense of status, ie prestige. It shows in their overwhelming mediocrity, eg the failure of economics to progress to a biologically scientific paradigm.
One reasonable approach might be to look at which group has produced the 'best' research over the past few years. But how do you judge that in a way that seems fair? Once you have a criteria to judge that, then people will start to game that criteria.
Or taking a step up, The university needs to save money. How do you judge if the Chemistry department or the Computer Science department should have its funding cut.
No matter how you slice it at some point you're going to need a way for someone to judge which of two departments is producing the 'best' research and thus deserves more money, and that will incentivize people to game that metric.
We aren't short on food, shelter, clothes, tech, etc - those are all solved problems.
The problem that isn't solved is stupid people sitting in charge of decisions they don't have the brain make-up to comprehend or manage, making pretend they know what they're doing, holding people far superior to them hostage.
Assuming not good faith for peer review would make academia more interesting, only way would probably for the peer reviewer go to the lab and get live measurements shown. Then check the equipment...
Still, reading your comment makes me despair. It plants a nagging doubt in my mind, "how many of these zillion studies cited that are actually replicable?" This doubt remains despite knowing that the scientist is one of the leading experts in the field, and very down-to-earth.
What are the solutions here? A big incentive-shift to reward replication more? Public shaming of misleading studies? Influential conferences giving more air-time for talks about "studies that did not replicate"? I know some of these happen at a smaller-scale[1], but I wonder about the "scaling" aspect (to use a very HN-esque term).
PS: Since I read Behave by Sapolsky — where he says "your prefrontal cortex [which plays critical role in cognition, emotional regulation, and control of impulsive behavior] doesn't come online until you are 24" — I tend to take all studies done on university campuses with students younger than 24 with a good spoon of salt. ;-)
It’s probably not all bullshit but I would bet a double digit percentage of it is.
Therefore it is not well suited to figure out the world.
You should treat all of it with extreme helpings of salt.
Also, I don't think ending poverty is a major stated goal of psychology research. . .
Many people might spare themselves at least some misery by educating themselves about evolutionary psychology, including the landmines and open questions.
I think the problem is much bigger than simply a binary is it replicable or not. It’s extremely easy to find papers by “leading experts” that have valid data with replicable results where the conclusions have been generalized beyond the experiments. The media does this more or less by default when reporting on scientific results, but researchers do it themselves to a huge degree, use very specific conditions and results to jump to a wider conclusion that is not actually supported by the results.
A high profile example of this is the “Dunning Kruger” effect; the data in paper did not show what the flowery narrative in the paper claimed to show, but there’s no reason to think they falsified the results. Some researchers have reproduced the results, as long as the conditions were very similar. Other researchers have tried to reproduce the results under different conditions that should have worked according to the paper’s narrative and conclusions, but found that they could not reproduce, because there were specific factors in the original experiment that were not discussed in the original paper’s conclusions -- in other words, Dunning and Kruger overstated what they measured such that the conclusion was not true. They both enjoyed successful academic careers and some degree of academic fame as a result of this paper that is technically reproducible but not generally true.
To make matters worse, the public has generally misinterpreted and misunderstood even the incorrect conclusions the authors stated, and turned it into something else. Almost never in discussions where the DK effect is invoked do people talk about the context or methodology of the experiments, or the people who participated in them.
This human tendency to tell a story and lose the context and details and specificity of the original evidence, the tendency to declare that one piece of evidence means there is a general truth, that is scarier to me than whether papers are replicable or not, because it casts doubt on all the replicable papers too.
There's obviously more complexity than this, but I believe that if even a relatively small percentage of the population started thinking like this (particularly, influential people) it could make a very big difference.
Unfortunately, this seems to be extremely counter to human nature and desires - people seem seem compelled to form conclusions, even when it is not necessary ("Do people have ideas, or do ideas have people?").
Out of curiosity, what's the title of the book?
[1] https://www.routledge.com/Evolutionary-Psychology-The-New-Sc...
Isn't that the moment where you try even harder to falsify the claims in that paper? You already know that you'll succeed so it wouldn't be a waste of time in your effort.
This is an example which did get cites:
https://journals.sagepub.com/doi/full/10.1111/j.1539-6053.20...
But despite the high visibility, you can see the large number of papers published based on the original myth.
And this refutation doesn't have great methodology (but other ones do). It's mostly cited due to strong language used.
The main problem is that even if you reproduce their experiment, they can claim that you did some step wrong, perhaps you are mixing it too fast or too slow, or the temperature is not correctly controlled, or that one of your reactive have a contamination that destroy the effect, or magically realize that their reactive that is important.
It's very difficult to publish papers with negative results. So there is a high chance it will not count in your total number of publications. Also, expect a low number of citation, so it's not useful for other metrics like citation count or h.
For the same reason, you will not see publications of exact replications. A good paper X will be followed by almost-replications by another teams, like "we changed this and got X with a 10% improvement" or "we mixed the methods of X and Y and unsurprisingly^W got X+Y". This is somewhat good because it shows that the initial result is robust enough to survive small modifications.
There are several levels of peer review. I've definitely been a reviwer on papers where the reviewers requested everything required and reproduced the experiment. That's extremely rare.
Some wider questions would be: Are there similar problems in Mathematics/physics versus the life sciences/other social sciences? Are there the same kind of problems across different fields of study?
Also i wonder if replication issues would be less severe if there was a requirement to publish the software and raw data that any study is based on as open source / data. It is possible that a change in this direction would make it more difficult to manipulate the results (after all it's the public who paid for the research, in most cases)
The only way to fix replication issues is to give financial and career incentives for doing replication work. Right now there are few carrots and many sticks.
Personally, I'm always very careful to cite and praise work by "competing" researchers even when that work has well-known errors, because I know that those researchers will review my paper and if there aren't other experts on the review committee the paper won't make it. I wish I didn't have to, but my supervisor wants to get tenured and I want to finish grad school, and for that we need to publish papers.
Lots of science is completely inaccessible for non-experts as a result of this sort of politics. There is no guarantee that the work you hear praised/cited in papers is actually any good; it may have been inserted just to appease someone.
I thought that this was something specific to my field, but apparently not. Leaves me very jaded about the scientific community.
I want to answer the question "if I were a researcher and were willing to cheat to get ahead, what should be the objective of my cheating?"
If you want to look impressive to non-experts and get lots of grant money/opportunities, I'd go for lots of straightforward publications in top-tier venues. Star findings will come under greater scrutiny.
If you want to write a pop book and on TV and sell classes, you need one interesting bit of pseudoscience and a dozen followup papers using the same bad methodology.
Does anyone have ideas on how that may be achieved - what a correct incentive structure for research might look like?
The trouble is that for the evaluators (all the institutions that can be sources of an incentive structure) it's impossible to distinguish an unpublished 90%-ready Nobel prize from unpublished 90%-ready bullshit. So if you've been working for 4 years on minor, incremental work and published a bunch of papers it's clear that you've done something useful, not extraordinary, but not bad; but if you've been working on a breakthrough and haven't achieved it, then there's simply no data to judge. Are you one step from major success? Or is that one step impossible and will never be achieved? Perhaps all of it is a dead end? Perhaps you're just slacking off on a direction that you know is a dead end, but it's the one thing you can do which brings you some money, so meh? Perhaps you're just crazy and it was definitely a worthless dead end? Perhaps everyone in the field thought that you're just crazy and this direction is worthless but they're actually wrong?
Peter Higgs was a relevant case - IIRC he said in one interview taht for quite some time "they" didn't know what to do with him as he wasn't producing anything much, and the things he had done earlier were either useless or Nobel prize worthy, but it was impossible to tell for many years after the fact. How the heck can an objective incentive structure take that into account? It's a minefield.
IMHO any effective solution has to scale back on accountability and measurability, and to some extent just give some funding to some people/teams with great potential, and see what they do - with the expectation that it's OK if it doesn't turn out, since otherwise they're forced to pick only safe topics that are certain to succeed and also certain to not achieve a breaktrhough. I believe European Research Foundation had a grant policy with similar principles, and I think that DARPA, at least originally, was like that.
But there's a strong entirely opposite pressure from key stakeholders holding the (usually government) purses, their interests are more towards avoiding bad PR for any project with seemingly wasted money, and that results in a push towards these broken incentive structures and mediocrity.
For commercial ventures, you also have the same issue of incremental progress vs big breakthroughs that don't look like much until they are ready.
As far as I can tell, in the startup ecosystem the whole thing works by different investors (various angels and VCs and public markets etc), all having their own process to (attempt to) solve this tension.
There's beauty in competition. And no taxpayer money is wasted here. (Yes, there are government grants for startups in many parts of the world, but that's a different issue from angels evaluating would-be companies.)
A 0.1% chance to build an app that's gonna be useful to hundreds of millions of people is better than what most career scientists manage.
You get what you measure for applies here. Now if we had some Objective Useful Research Quality Score t could replace the price signals. But then we wouldn't have the problem in the first place, just promote based on OURQS.
At the same time, academics have been increasingly been evaluated by some metrics to show value for money. This has let to some schizophrenic incentive structures. Most professor level academics are spending probably around 30% of their time on writing grants, evaluating grants and reporting on grants. Moreover, the evaluation criteria also often demand that work should be innovative, "high risk/high reward" and "breakthrough science", but at the same time feasible (and often you should show preliminary work), which I would argue is a contradiction. This naturally leads to academics overselling their results. Even more so because you are also supposed to show impact.
The main reason for all this IMO is the reduced funding for academic research in particular considering the number of academics that are around. So everyone is competing for a small pot, which makes those that play to the (broken) incentives, the most successful.
> the goal of research, which is to expand the scope and quality of human knowledge.
But are we so certain this is ever what drove science? Before we dive into twiddling knobs with a presumption of understanding some foundational motivation, it's worth asking. Sometimes the stories we tell are not the stories that drive the underlying machinery.
For e.g., we have a lot of wishy-washy "folk theories" of how democracy works, but actual political scientists know that most of the ones people "think" drive democracy, are actually just a bullshit story. According to some, it's even possible that the function of these common-belief fabrications is that their falsely simple narrative stabilizes democracy itself in the mind of the everyman, due to the trustworthiness of seemingly simple things. So it's an important falsehood to have in the meme pool. But the real forces that make democracy work are either (a) quite complex and obscure, or even (b) as-of-yet inconclusive. [1]
I wonder if science has some similar vibes: folks theory vs what actually drives it. Maybe the folk theory is "expand human knowledge", but the true machinery is and always has been a complex concoction of human ego, corruption and the fancies of the wealthy, topped with an icing of natural human curiosity.
[1]" https://www.amazon.ca/Democracy-Realists-Elections-Responsiv...
The Structure of Scientific Revolutions by Thomas Kuhn is an excellent read on this topic - dense but considered one of the most important works in the philosophy of science. It popularized Planck's Principle paraphrased as "Science progresses one funeral at a time." As you note, the true machinery is a very complicated mix of human factors and actual science.
Perhaps start with removing tax payer money from the system.
Stop throwing good money after bad.
"Academic politics is the most vicious and bitter form of politics, because the stakes are so low."
As a non-expert, this is not the type of inaccessibility that is relevant to my interests.
"Unfortunately, alumni do not have access to our online journal subscriptions and databases because of licensing restrictions. We usually advise alumni to request items through interlibrary loan at their home institution/public library. In addition, under normal circumstances, you would be able to come in to the library and access the article."
This may not be technically completely inaccessible. But it is a significant "chilling effect" for someone who wants to read on a subject.
Also, I know of no researchers personally who are enthralled by the existing system.
If I am reading between the lines correctly, you are implying there are few undergrads publishing in high caliber journals because of gatekeeping. As a reviewer, I often don't even know the authors' names, let alone their degrees and affiliations. It is theoretically possible that editors would desk reject undergrads' papers, but: a) I personally don't think a PhD is required to do quality research, especially in CS, and I know I am not the only person thinking that; b) In some fields like psychology and, perhaps, physics many junior PhD students only have BS degrees, which doesn't stop them from publishing.
I think that single-authored research papers by people without a PhD are relatively uncommon because getting a PhD is a very popular way of leveling up to the required expertise threshold and getting research funding without one is very difficult. I don't suspect folks without a PhD are systematically discriminated against by editors and reviewers, but, of course, I can't guarantee that this universally true across all research communities.
I think the entire academic enterprise needs to be burnt down and rebuilt. It’s rotten to the core and the people who are providing the most value - the scholars - are simultaneously underpaid and beholden to a deranged publishing process that is a rat race that accomplishes little and hurts society. Not just in our checkbook but also in the wasted talent.
It also isn't any sort of conspiracy that government grants are given out to people with a proven history of doing good research, as evaluated by their peers.
The people you mention are probably making YouTube videos and writing blog posts about their findings and are reaching a broader audience..
"Neither Myers nor Briggs was formally educated in the discipline of psychology, and both were self-taught in the field of psychometric testing."
Maybe not quite as prestigious as nature, but NLP is pretty huge and the conference I got into has average h index of I think 60+
BTW I can tell you the the vast majority of researchers are not "enthralled" by the system, but highly critical. They simply don't have a choice but to work with it.
I think the inaccessibility is for different reasons, most of which revolve around the use of jargon.
In my experience, the situation is not so bad. It is obvious who the good scientist are and you can almost always be sure that if they wrote it it's good.
In essence, the evaluators (non-scientific organizations who fund scientific organizations) need some metric to compare and distinguish decent research from weak, one that's (a) comparable across fields of science; (b) verifiable by people outside that field (so you can compare across subfields); (c) not trivially changeable by the funded institutions themselves; (d) describable in an objective manner so that you can write up the exact criteria/metrics in a legal act or contract. There are NO reasonable metrics that fit these criteria; international peer-reviewed-publications fitting certain criteria are bad but perhaps least bad from the (even worse) alternatives like direct evaluation by government committees.
(I am leaving cetacean cunt in because it’s a funny autocorrect.)
(And now I’m leaving the above in, because it’s even funnier. Both genuine.)
(There's lots of crowding out happening, of course, from the government subsidized science. But that can't be helped at the moment.)
It will probably have to be started by some civic minded billionaires. I don't think the established system can reform itself.
I wouldn’t necessarily condone the behavior, but what would you do in the situation? To always whistleblow whenever something doesn’t feel right and risk the politics? To quit working in the field if your concerns aren’t heard? To never cite papers that have absolutely any errors? I think it’s a tough situation and not productive to say OP isn’t behaving morally.
The situation is even worse when the paper claiming X underwent artifact review, where reviewers actually DID look at the raw data and source code but simply lacked the attention or expertise to recognize errors.
I'm not taking bribes, I'm paying a toll.
What you've described sounds like something that is not, in any sense, science.
From your perspective, what can be done to return the scientific method to the forefront of these proceedings?
But one other thing to note here is that these headlines about a "replication crisis" seems to imply that this is a new phenomenon. Let's not forget the history of the electron charge. As Feynman said:
"We have learned a lot from experience about how to handle some of the ways we fool ourselves. One example: Millikan measured the charge on an electron by an experiment with falling oil drops, and got an answer which we now know not to be quite right. It's a little bit off because he had the incorrect value for the viscosity of air. It's interesting to look at the history of measurements of the charge of an electron, after Millikan. If you plot them as a function of time, you find that one is a little bit bigger than Millikan's, and the next one's a little bit bigger than that, and the next one's a little bit bigger than that, until finally they settle down to a number which is higher. Why didn't they discover the new number was higher right away? It's a thing that scientists are ashamed of—this history—because it's apparent that people did things like this: When they got a number that was too high above Millikan's, they thought something must be wrong—and they would look for and find a reason why something might be wrong. When they got a number close to Millikan's value they didn't look so hard. And so they eliminated the numbers that were too far off, and did other things like that ..."
https://en.wikipedia.org/wiki/Oil_drop_experiment#Millikan.2...
https://en.wikipedia.org/wiki/Cosmic_ray_visual_phenomena
It's interesting that according to the Wikipedia article it's not entirely certain whether the radiation is producing actual light or just the sensation of light.
The social sciences face the problem of not having so many different possible angles, such as quantitative theories or even a clear idea of what is being tested. Much of the research is engaged in the collection of isolated factoids. Hopefully something like a quantitative theory will emerge, that allows these results to be connected together like a mesh network, but no new science gets there right away.
The other thing is, to be fair, social sciences have to deal with noisy data, and with ethics. There were things I could do to atoms in my experiments, such as deprive them of air and smash them to bits, that would not pass ethical review if performed on humans. ;-)
PBS Space time has an excellent video on the topic: https://www.youtube.com/watch?v=72cM_E6bsOs
In physics, when a result raises more questions than it answers, we call it "job security." ;-)
No, it isn't. It looked at a few different fields, and found that the problem was actually worse for general science papers published in Nature/Science, where non-reproducible papers were cited 300 times more often as reproducible ones.
I've seen papers without any sort of value or reason to exist being bruteforced through reviewing just to avoid some useless junk data of no value whatsoever being wasted, all to just add a line on someone's CV.
This is without saying that some Unis are packed of totally incompetent people that only got to advance their careers by always finding a way to piggyback on someone else's paper.
The worst thing I've seen is that reviewing papers is also often offloaded to newly graduated fellows, which are often instructed to be lenient when reviewing papers coming from "friendly universities".
The level of most papers I have had the disgrace to read is so bad it made me want to quit that world as soon as I could.
I got to the conclusion the whole system is basically a complex game of politics and strategy, fed by a loop in which bad research gets published on mediocre outlets, which then get a financial return by publishing them. This bad published research is then used to justify further money being spent on low quality rubbish work, and the cycle continues.
Sometimes you get to review papers that are so comically bad and low effort they almost feel insulting on a personal level.
For instance, I had to reject multiple papers not only due their complete lack of content, but also because their English was so horrendous they were basically unintelligible.
Maybe this is how the GPT "AI" can generate such similar results. Lol.
.... but each field is different. For those that are more quantitative, it's harder to deviate you conclusion from the data.
Bias is not binary, so it's a sliding scale between the hard sciences and the squishy ones.
You have to listen to the science, and also use the common sense that "this is as far as we know" and that knowledge today may change tomorrow.
Two comments below you use this "argument" to ask "evidence" for evolution and climate change. Big red flag.
If you think there is no incorrect science at all in regards to evolution and climate change you're no better than the zealots of any religion.
The government can then reduce funding to institutions that have too high a percentage of research that failed to be replicated.
From that point the situation should resolve itself as institutions wouldn’t want to lose funding - so they’d either have an internal group replicate before publishing or coordinate with other institutions pre-publish.
Anything I’m missing?
That's no guarantee; in 2020 the US essential oils industry was worth $18.62 billion: https://www.grandviewresearch.com/industry-analysis/essentia... . Which is bigger than the US music recording industry (https://www.musicbusinessworldwide.com/the-us-recorded-music...).
- You would need some sort of barrier preventing movement of researchers between these audit teams and the institutions they are supposed to audit otherwise there would be a perverse incentive for a researcher to provide favorable treatment to certain institutions in exchange for a guaranteed position at said institutions later on. You could have an internal audit team audit the audit team, but you quickly run into an infinitely recursive structure and we'd have to question whether there would even be sufficient resources to support anything more than the initial team to begin with.
- From my admittedly limited experience as an economics research assistant in undergrad, I understood replication studies to be considered low-value projects that are barely worth listing on a CV for a tenure-track academic. That in conjunction with the aforementioned movement barrier would make such an auditing researcher position a career dead-end, which would then raise the question of which researchers would be willing to take on this role (though to be fair there would still be someone given the insane ratio of candidates in academia to available positions). The uncomfortable truth is that most researchers would likely jump at other opportunities if they are able to and this position would be a last resort for those who aren't able to land a gig elsewhere. I wouldn't doubt the ability of this pool of candidates to still perform quality work, but if some of them have an axe to grind (e.g. denied tenure, criticized in a peer review) that is another source of bias to be wary of as they are effectively being granted the leverage to cut off the lifeline for their rivals.
- You could implement a sort of academic jury duty to randomly select the members of this team to address the issues in the last point, which might be an interesting structure to consider further. I could still see conflict-of-interest issues being present especially if the panel members are actively involved in the field of research (and from what I've seen of academia, it's a bunch of high-intellect individuals playing by high school social rules lol) but it would at least address the incentive issue of self-selection. Perhaps some sort of election structure like this (https://en.wikipedia.org/wiki/Doge_of_Venice#:~:text=Thirty%....) could be used to filter out conflict of interest, but it would make selecting the panel a much more involved and time-consuming process.
The desire to control and incentivize researchers to compete against each other in order to justify their salary is understandable, but it looks like it has been blown so out of proportions lately that it's doing active harm. Most researchers start their career pretty self-motivated to do good research.
Installing another system to double-check every contribution will just increase the pressure to game the system in addition to doing research. And replicating a paper may sometimes cost as much as the original research, and it's not clear when to stop trying. How much collaboration with the original authors are you supposed to do, if you fail to replicate? If you are making decisions about their career, you will need some system to ensure it's not arbitrary, etc.
When people deliberately fake lab data to further their career, and that fake data is used to perform clinical trials on actual people, that's not just fraudulent, it's morally destitute. Yet this has happened.
People deliberately use improper statistics all the time to make their data "significant". It's outright fraud.
I've seen people doing sloppy work in the lab, and when questioning them, was told "no one cares so long as it's publishable". Coming from industry, where quality, accuracy and precision are paramount, I found the attitude shocking and repugnant. People should take pride and care in their work. If they can't do that, they shouldn't be working in the field.
PIs don't care so long as things are publishable. They live in wilful ignorance. Unless they are forced to investigate, it's easiest not to ask any questions and get unpleasant answers back. Many of them would be shocked if they saw the quality of work done by their underlings, but they live in an office and rarely get directly involved.
I've since gone back to industry. Academia is fundamentally broken.
When you say "double-checking" won't solve anything, I'd like to propose a different way of thinking about this:
* lab notebooks are supposed to be kept as a permanent record, checked and signed off. This rarely happens. It should be the responsibility of a manager to check and sign off every page, and question any changes or discrepancies.
* lab work needs independent validation, and lab workers should be able to prove their competence to perform tasks accurately and reproducibly; in industry labs do things like sending samples to reference labs, and receiving unknown samples to test, and these are used to calculate any deviation from the real value both between the reference lab and others in the same industry. They get ranked based upon their real-world performance.
* random external audits to check everything, record keeping, facilities, materials, data, working practices, with penalties for noncompliance.
Now, academic research is not the same as industry, but the point I'm making here is that what's largely missing here is oversight. By and large, there isn't any. But putting it in place would fix most of the problems, because most of the problems only exist because they are permitted to flourish in the absence of oversight. That's a failure of management in academia, globally. PIs aren't good managers. PIs see management in terms of academic prestige, and expanding their research group empires, but they are incompetent at it. They have zero training, little desire to do it, and it could be made a separate position in a department. Stop PIs managing, let them focus on science, and have a professional do it. And have compliance with oversight and work quality part of staff performance metrics, above publication quantity.
However, I think there is potential in taking the 'funded by the government' idea in a different direction. Having a publication house that was considered a public service, with scientists (and others) employed by the government and working to review and publish research without commercial pressures could be a way to redirect the incentives in science.
Of course this would be expensive and probably difficult to justify politically, but a country/bloc that succeeded in such long term support for science might end up with a very healthy scientific sector.
Thank goodness our newsmedia business doesn't work that way, or we would be poorly-informed in multiple ways.
> Prediction markets, in which experts in the field bet on the replication results before the replication studies, showed that experts could predict well which findings would replicate (11).
So it's even stating that this isn't completely innocent, given different incentives most reviewers identify a suspicious study, but under current incentives it seems letting it through due to the novelty is somehow warranted.
People love this stuff. Malcolm Gladwell's made a career on it: half of the stuff he writes about is disproven before he publishes. It's very interesting that facial microexpressions analysis can predict relationship outcome with 90% certainty. Except it's just an overfit model, it can't, and he's no longer my favorite author. [0]
Similarly, Thomas Erikson's "Surrounded by Idiots" also lacks validation. [1]
Both authors have been making top 10 lists for years, and Audible's top selling list just reminded me of them.
Similarly, shocking publications in Nature or Science are to be viewed with skepticism.
I don't know what I can read anymore. It's the same with politics. The truth is morally ambiguous, time consuming, complicated, and doesn't sell. I feel powerless against market forces.
[0] https://en.wikipedia.org/wiki/John_Gottman#Critiques
[1] https://webcache.googleusercontent.com/search?q=cache:5Z7JiC...
But here's the thing, people don't have time for this. They have work, bills, home and car maintenance, groceries, kids, friends, a slew of media to consume, recreation on top of it - and they're all dying. So it doesn't matter who says what, they're going to pick the dilute politicized version of the results that their team supports and run with it regardless of what the nigh-unreadable highly specialized papers say. Orwell said it well "I believe that this instinct to perpetuate useless work is, at bottom, simply fear of the mob. The mob (the thought runs) are such low animals that they would be dangerous if they had leisure; it is safer to keep them too busy to think."
And those who do elect the burden of extracurricular mental activity aren't given much in the way of options in any case. What are they to do, disseminate the material to their friends, co-workers, children - quite probably the very same population as mentioned above, weighted with the ceaseless demands of reality? To what end? Chinese whispers? It's better to have them say, "I don't know, I'm not convinced either way." A construal which is developed from adequately exercised critical skills. But that's another discussion about perverse social conditioning no doubt evolved from the deployment of poorly understood technique compounded by its acceptance as custom in education - I'm speaking of course about grading and student assessment. Nobody wants to be stupid at the very least, and professing one's ignorance is construed as an admission of guilt.
It's called the "research grant system"
I don’t know what the answer is, and I’ve been worried for a while that we are putting blind faith in “science” which just lines up with our preferred worldview. Maybe the answer is simply to use science, however it is performed, to inform, not guide, policy, and always keep in mind that what science believes has a non-zero chance of being politically-driven itself.
"Our findings directly contradict [1]" is a citation.
Without context, number of citations doesn't tell you anything.
https://fantasticanachronism.com/2020/09/11/whats-wrong-with...
"You might hypothesize that the citations of non-replicating papers are negative, but negative citations are extremely rare.5 One study puts the rate at 2.4%. Astonishingly, even after retraction the vast majority of citations are positive, and those positive citations continue for decades after retraction.6"
Is this really the case? And is this actually a "new" phenomenon?
It seems like it could be a disguised version of the Availability Cascade. [1] In other words, when we encounter a simple-to-understand explanation of something complex, the explanation ends up catching on.
Then, because the explanation is simple, its popularity snowballs. The idea cascades like a waterfall throughout the public. Soon it becomes common sense—not because of sense, but because of common.
The number of academic “big shots” (friends of the poster, not of me) who “liked” the comment was a bit alarming.
There’s too much incentive for fudging things (depending on your field, either grants or company funding).
The degree of fraud in Chinese journals is high and well discussed (as it should be). But apart from a small amount of hand-wringing over “the replication crisis” there is no similar condemnation of the work in the rest of the world.
It is also quite telling that the biggest differences in citation counts are for papers published in Nature and Science. But in discipline specific journals (Figs. 1B,C), the effect is very modest. Practicing scientists know that Science and Nature are publish the least reproducible results, in part because they like "sexy" (surprising, less likely to be correct) science, and in part because they provide almost no detail on how experiments were performed (no Materials and Methods).
The implication of the paper is that less reproducible science has more impact than reproducible science. But we know this is wrong -- reproducible results persist, while incorrect results do not (we haven't heard much more about the organisms that use arsenic rather than phosphorus in their DNA -- https://science.sciencemag.org/content/332/6034/1163 )
The phenomenon described in the articles sounds like a natural consequence of this attitude.
The incentives for Nature are not to produce great science, but to sell journals and that requires them to give the impression of being on the forefront of "scientific discovery". I've in fact been told by an editor "our job is to make money, not great science)".
The irony is that their incentives also make them very risk averse, so they will not publish results which don't have a buzz around them. I know of several papers which created new "fields" which were rejected by the editors. The incentive also results in highly cited authors having an easier time to get published in Nature.
I should say that this is typically much better in the journals run by expert editors, published by the technical societies like e.g. IEEE.
Obviously, because journals don't attempt replication, nor should they.
A study will only have replication by another researcher/team after it's been published. Or not, in which case that's publishable as well.
This is how science works and is supposed to work.
The problem should rather be: why are journals accepting papers with citations to debunked papers that don't also cite the debunking papers?
I've had plenty of friends have a paper get sent back for revisions because it was missing a newer citation. This should be a major responsibility of reviewers. So why aren't peer reviewers staying on top of the literature?
Perhaps because reviewing is unpaid work, both in the financial and academic sense?
This is a major issue with the current model.
To me it seems all our social structures are decaying. Even look at the language and how everything is faux hyperbole these days. People are forgetting how to think, how to write, how to reason all for the sake of the woke cultural revolution. The political correctness of thought is valued over the ontological correctness.
Have some portion of academic funding distributed by assigning a pool of money to grantors and having them bet portions of that money on papers they consider impactful AND likely to replicate. Allow the authors to use the money on any research they want.
This lets authors get some research money without writing costly and wasteful grant proposals. It takes advantage of the fact that experts in the field can generally tell which studies are likely to replicate (I'm assuming granting agencies can find experts in the field to do this).
> At a global scale, anything that can happen will happen a small but nonzero times: this has been epitomized as “Littlewood’s Law: in the course of any normal person’s life, miracles happen at a rate of roughly one per month.” This must now be extended to a global scale for a hyper-networked global media covering anomalies from 8 billion people—all coincidences, hoaxes, mental illnesses, psychological oddities, extremes of continuums, mistakes, misunderstandings, terrorism, unexplained phenomena etc. Hence, there will be enough ‘miracles’ that all media coverage of events can potentially be composed of nothing but extreme outliers, even though it would seem like an ‘extraordinary’ claim to say that all media-reported events may be flukes.
The first Tesla Model S isn't anywhere as good as the current version, but it was still groundbreaking. Academic publications work the same way.
The problem is that some fields don’t do replication studies. In such fields these highly cited papers with dubious findings are never debunked, because you know economists to pick on one specific group, don’t do replication studies.
1/ For best journals: a non-replication betting market for peer-reviewed and published papers. Grants to replicate papers that are highest in the betting pool.
2/ Citation index were citing and publishing a paper that does not replicate lowers the score.
(This should be first used in machine learning research)
To add insult to injury: when subsequent experiments didn't give any interesting results, and the PhD didn't have enough material for graduation, they decided to split the data they already had by gene variants, because technology had just reached the point where that kind of analysis could be done by a sufficiently rich lab. And there was a new result! The original blob was now seen in two areas instead of one, depending on variant. Highly publishable. That that meant that the original finding was now invalidated wasn't even considered.
edit: this took place between 2000 and 2010.
Take for example a study done in about 2005 that said "willpower is an exhaustible resource" - They gave some participants some problems to solve - they were cooking cookies at that back so the participants could smell them. Later, half of the participants were given cucumbers and the other half cookies and then observed that people given cookies persisted for longer in solving the problems. This proved "willpower is an exhaustible resources".
Is willpower exhaustible? what if the participants while they were giving up early where given an incentive of 500 USD to continue working on the same problem, would they work for much longer than the other group? why? willpower is an exhaustible resource, how can you draw from something that now no longer exists? would you then conclude that willpower can be depleted in humans, but humans have the ability to synthesize it from currency. Ofcorse thats ridiculous.
Willpower and expended effort are context based and each person can set their own threshold of when its appropriate to give up - its partly a conscious decision. A few years later the authors of this same study walked back on their claims that "willpower is an exhaustible resource" ( it does not stop people from touting this around as reasoning for whatever point they are trying to make ).
My point is bio-monitoring technology with probably VR will atleast give us some verifiable data-points..
As an example: in the field in which I worked, it was commonly assumed that the processing mechanism had to interpret a part of the stimulus as belonging to one thing or the other. Then they would manipulate the context to see how they could influence that. There were many results and experimental refinements, leading to an endless stream of contradictory articles, all claiming victory for their theory. But there are models in which this assumption doesn't make sense. Instead, it could belong to both, and the ambiguity would be resolved later. That makes the whole research line I described nonsense. The data is valid (if well documented), but the analyses and conclusions are not.
A genetics statistician explained to me that all early genetic research (and that means up until 2000 at least) is highly unreliable, because they used Neyman-Fisher significance testing with high p values, like 0.01. The lack of understanding was (and probably still is) so great, that 0.01 was simply not enough to exclude random results. Nearly all those papers are irreproducible. Nowadays, 5 sigma, which is less than 0.000001, is considered safe, although he himself was convinced that Bayesian is the only way to go.
The understanding of the subject in psychology is considerably worse than in biology. No amount of bio-monitoring is going to overcome that.
The amount of interest from regular people and in the subject and the demand for research is so much that people want to hang on to twigs.. most people want to improve themselves and be able to think and do better and take atleast a fleeting interest in psychology... Evidenced by the rise of Jordon Peterson -
I recently found Dr.Andrew Huberman to be a voice of reason (check out his Youtube podcast, its very information-dense and pretty good) - He approaches things from a neurobiology perspective. To navigate around getting a deeper understanding yourself, I thing a chemical / hormone / neurotransmitter based discussion at least grounds your feet in real things, from which you can probably explore things for yourself.. It also eliminates a lot of the cruft from the discussion..
https://cfr.pub/forthcoming/chang-li-2018.pdf
The experience ucompletely unraveled my faith in academia.
The real issue is that citation count is a terrible metric and always has been (when I was doing research the magic you wanted was the low but not no citation count paper which went into a huge amount of detail on exactly what they did).
Replication failures disproportionately hit high-profile journals. The journals' editorial processes favor "surprising", counter-intuitive claims: life based on arsenic, stem cells created by treating skin cells with a mild acid, cold fusion, etc. The prior probability these are true is lower (hence the surprise). At the same time, this also makes it more likely that someone will attempt to replicate the results--and that there will be interest in a negative paper if they can't.
Maybe that was reasonable given how science was advancing at the time. If so, what powers and advances have poor research systems denied us?
Alternatively, it may be that all the low hanging fruit in the known orchards are picked. But there are still workers reaching for the remaining higher, rarer fruit. Only a few are needed for what fruit remains. But there is instead a glut of workers. A few are up to the task. The rest pick up leaves and make convoluted arguments about their potential value. "These leaves are edible!"
Maybe we need can find new orchards. Or maybe the unpicked orchards are too far away to ever reach. Or maybe they don't exist at all.
But in most marketplaces advertising is a dominant force, and advertising power is not a function of product quality; you can market a bad product successfully. There are standards of truth in advertising, but they're loose and poorly enforced, the burden of evaluation falls on the consumer. Additionally, tribalist behavior builds up around products to a certain extent, as buyers of product X who dislike hearing it's good or crap push back on such claims, for varying reasons. Where these differences are purely aesthetic that doesn't matter much, but where they're functional a objectively better product can lose out against a one from a dishonest competitor or one with an irrationally loyal following. A problem for both producers and consumers is that it's more expensive to refute a false claim than to make it, so bad actors are incentivized to lie and lock int he advantage; it's arguably cheaper to apologize if caught out than to forgo the profitable behavior, as exemplified in the aphorism 'it's easier to ask forgiveness than permission.'
We see the results with problems like the OP and also in things like debates about public health, vaccine safety, climate change, and many other political issues. It's profitable to lie, a significant number of people have no problem doing it, the techniques of doing so have been repeatedly refined and weaponized, and those who rely on or simply prefer truth end up at a significant economic disadvantage. don't make the mistake of thinking this is problem confined to the niche world of academic journals.
References are often ambiguous. A mention like “Although some have argued X[1]” could be skeptical, but not explicit, about Reference #1’s quality. I could mean that there’s convincing data on both sides, or I could mean that Ref #1 is hot garbage (but don’t want a fight).
In some cases, there might be legitimately useful information in a retracted paper. If a paper describe an experiment, but then shows faked results, I’m not sure it’s wrong to cite it if you use a similar setup; that is where the idea came from, after all.
Most critically, there’s an unpredictable and often long lag between reading a paper, citing it one’s own work, and that work being published. I’ve had things sit at a journal for a year before publication, and it never occurred to me to “re-verify” the citations I included; indeed, I’ve never heard of anyone doing that.
We haven't scaled the practice of science for the 21st century.