Data detectives spotted fake numbers in a widely cited paper
economist.com
economist.com
In particular, this is not the first such issue for Dan Ariely, as Gelman points out, he has a history of sketchy scientific ethics like doing media tours for studies that he knows failed to replicate.
[0] http://datacolada.org/98 [1] https://statmodeling.stat.columbia.edu/2021/08/19/a-scandal-...
It feels like this could be applied retroactively too.
In this case the 2012 authors still had the data that they released in 2020 which is how the analysis got done that showed evidence of fraud. Might be worth just asking a whole bunch of people to release data they previously hadn't and collectively putting some time and effort into that.
For what it’s worth, the Trump administration attempted to make issuing new health and environmental regs harder by requiring public data disclosure. They did this entirely because they knew that much of the data could not be disclosed. So if you were studying, eg, the effects of some pollutant on a health outcome using private data, you wouldn’t be able to rely on that study in a regulatory context bc the data could not be published.
It’s a worthy idea, but there are exceptions for good reasons.
Annoyingly, the NPR transcript at [1] only has a small note "(SOUNDBITE OF ARCHIVED NPR BROADCAST)" at the top with no indication of when (the audio doesn't seem to have a date either). The podcast show notes are apparently the only recordation of the date. [2]
[1] https://www.npr.org/transcripts/805808486
[2] https://pbs.twimg.com/media/E9UQv8LWEAkpoR0?format=jpg&name=...
You start with a sexy story that you know will get you a lot of press, like "promising you will be honest actually makes you behave in an honest way" and then you just make that paper happen, however you can.
Under the publish or perish system, scientists don't have time to actually research the topic, and imagine if it fails to confirm - you just wasted a lot of time and didn't publish anything. Too risky, it's much easier to just fake it till you make it, especially since you know peer reviewers never ever will accuse you of fraud.
Any reviewer accusing a scientist of fraud will just be excluded from the community, since it's very important to uphold the narrative that "scientists are always honest, they never cheat like politicians, which is why we must always trust scientists and never question them".
https://www.discovermagazine.com/planet-earth/the-scientific...
Though i will say that they missed Robert Hooke's essay on the scientific method. I swear no one knows about this (even though Hooke was a founding and seminal member of the English Royal Society) because Hooke sounds insane. Who makes titles like this? I love it:
A scheme, or idea of the present state of natural philosophy, and how its defects may be remedied by a methodical proceeding in the making experiments and collecting observations wereby to compile a natural history, as to the solid basis for the superstructure of philosophy
https://plato.stanford.edu/entries/scientific-method/
And I think the demarcation problem can be in some ways be solved by distinguishing on predictivity.
Reading reports like this, I can see why there has been so much criticism in recent years of the concept of the banality of evil, and why so many researchers affect concern for the environment, because more than anyone, they seem to understand what it means to be responsible for poisoning an ecosystem and they need to get out in front of those narratives. Sure, it's just a bit of fraud in a journal, just like it's just a bit of PCB or mercury in a lake, and only a minority of the population who will be impacted, but someone has to call it out as nihilism, or we're a party to it as well.
As much as I dislike blockchain ideas, it makes me think a blockchain DAG of metadata about the integrity of published papers for citations, reproducability, and evidence of certain types of fraud would rebalance the incentives a bit.
Lets say there are 200 propositions we want to test and are candidates for publication, that 20 of them are true and that our error rate is 5%. That means when we test the 20 that are true 19 of them will be accurately shown to be correct and 1 will be erroneously found false.
However when the other 180 propositions are tested 5% of them will be erroneously found to be true, that's 9 propositions. This means we will end up with 28 'successful' studies that make it into prestigious journals, about 1/3 of which are false positives.
And as I said, that's if the system works perfectly with no fraud whatsoever. Throw in some human error and it's not surprising if a fair few studies start to look pretty dodgy. Add in some genuine fraud too and you've got a full-on replication crisis with all the trimmings.
The well-documented failings of psychology should cause us to take a closer look at other fields to see if they have similar problems. If they don't, then great. If they do, then fix them.
For example people studying metascience have found that a lot of medical research is of questionable accuracy. It is not as bad as psychology. But it is bad and I'm glad that people are taking this problem seriously.
I came at it from another direction here — there are people who are like: "Look at the replication crisis in psychology — this is proof science cannot be trusted in general". So what I meant to argue here is that this conclusion cannot be drawn automatically, not that we shouldn't scrutinize other fields (we should!)
Gigantic accusation, zero evidence.
If this was a paper, rather than an HN comment, I'd say there was every chance that it would be self-illustrating.
Any author of a published paper who will not stand behind the paper should have their name removed.
When there is just one name left, that person either accepts responsibility for the content, or they too disavow it and get removed.
When there are no names left, the paper is retracted.
The spam in these journals puts Buzzfeed to shame.
A study on dishonesty was based on fraudulent data - https://news.ycombinator.com/item?id=28271805 - Aug 2021 (42 comments)
Noted study in psychology fails to replicate, crumbles with evidence of fraud - https://news.ycombinator.com/item?id=28264097 - Aug 2021 (102 comments)
A Big Study About Honesty Turns Out to Be Based on Fake Data - https://news.ycombinator.com/item?id=28257860 - Aug 2021 (90 comments)
Evidence of fraud in an influential field experiment about dishonesty - https://news.ycombinator.com/item?id=28210642 - Aug 2021 (51 comments)
And no one notices. (or don't care about the obviously suspicious aspect)
Do scientists actually read what they cite?
I depends on ones probity, but yes the phenomenon is widespread. There are numerous reasons to cite a paper and just read it superficially: padding the references list, pleasing a reviewer by adding a paper he recommended, citing friends, etc. In my own lab they frequently cite a certain theory which if you actually read about it has nothing to do with what they are doing.
No. What are you kidding me. There's like typically over 50 papers referenced in a typical publication, no way I have the time to read all of them carefully
Yes, but few actually scrutinize the methodology of the studies. Statistics is really hard. It's easier to assume peer reviewers would have rejected the paper if it was bad.
The AUTHORS of the original paper got a dataset from a company. They didn't assume the fraud from the start and published the paper based on it.
Later, when they tried to analyze the issue more in-depth, they couldn't replicate the results. THE ORIGINAL AUTHORS PUBLISHED a paper about a failure to replicate. It was just then that someone looked at the original data and found that it was faked.
The fear of reputational implosion is apparently insufficient.
I did completely read what was available to me without having an account.
> The AUTHORS of the original paper got a dataset from a company. They didn't assume the fraud from the start and published the paper based on it.
My comment is not about who the culprit is or isn't. Indeed, I don't mention anything about it.
Rather, it's about how, as the title says, a WIDELY cited paper has fabricated data following rather (IMO) obvious red flag patterns and none of the people -who cited the paper- raised issues about that.
Thus, I questioned whether scientists read or not the papers they cite in the parent post. The question is not a judgment, I'm just truly curious since I'm not part of the formal academia, just an undergraduate.