What happens when scientists admit error
elemental.medium.com
elemental.medium.com
https://sites.tufts.edu/histmath/files/2015/11/Frege-Letter-...
Russell himself says "As I think about acts of integrity and grace, there is nothing to compare with Frege's dedication for truth."
This is why fake news is such a hit ;)
Yet it is hard to believe scientists would be the ones who only ship flawless software. And easier to imagine that most often, when one realises a previous publication's result was indeed impacted by an error and faces the difficult decision to retract or pretend they haven't noticed, they opt for the other route. I've found scientists in general to have less ego then average, as expected from people trained to care only for the facts, but they still operate in the larger framework of individualistic modern society. It probably takes a scientist and a woman, whose education generally encourage lesser ego to begin with, to admit such an error.
Or a software author, but there there is no other alternative to guilt admittance than total ridicule, since we are drowning everyday in a world of errors.
I mean, most of them don't even pin down their dependencies. I don't want to touch a lot of python code in the open written by scientists.
Thinking about software reproducibility is the last thing I have noticed on repositories submitted with a paper on arxiv and other places of publishing.
Not even a requirements.txt or mentioning what they used.
1. Reproducibility
This is very important for scientists so I expect them to know about package management and containers. You don't need to know everything about it.
Virtual environments, docker and using a good package manager like poetry (which is just an abstraction over virtualenv) shouldn't take more than a few weeks. Once you are done with the basics, you can learn more as you build.
2. Git and documentation
Add another few weeks.
They don't need to write good code. They just need to document it which they should already be good at. Optimization, scalability, distribution and other things can be handled by engineers but without documentation, it makes the work a lot harder.
This seems to be true if the code is the product of one person working for a few weeks or 20,000 people who work for 10 years.
A few years ago I had a conversation where the team insisted to me that the extensive documentation and instructions for a package were worthless because "good code is self documenting" the leader of the pack had shiny eyes and banged the table several times.
Be cautious, be kind and look for value in the code that you are using - not just problems. Also, consider: if you can pick up a code base (written by someone who knows the science) would it not be more profitable to repair, document and improve rather than demanding that the scientists learn both your skills and your way of doing things? If you conclude the latter, would you (hand on heart) be able to empirically demonstrate that your way of doing things is the right way?
Reproducability is a check that you got the right answer. As you describe it, it's more like a crappy programmer that relies on the testing team to catch his errors rather than be responsible for his own.
> and other things can be handled by engineers
Due to lack of funding and low pay I think most of them don't have access to software domain experts (aka seasoned programmers)
In practice most grad students get coding experience by working on projects, which is a great way to learn, but its rare for anyone to look at the code itself afterward if the results are sensible. I dont know of any small group with a formal code review process. I think we could make relatively inexpensive improvements in science by bringing in more dedicated software people who can educate others and help with the software engineering aspects. Its not unusual for a lab with half a dozen groups and 50+ scientists to have no software specialists, and 1 or 2 IT people whose responsibilities are purely infrastructure oriented.
I have met a few "born SREs" who had little or no training but knew exactly what to do in production outage situations. But I haven't met anybody who was a "Great Software Engineer" out of the box.
And scientists aren't going to take a lot of time out of their education to get that day-to-day programming experience. I was lucky- I worked in a theory lab, next to a team of good software developers, and worked/learned from them over a period of years.
I think what we need are more research engineers, who are people who do software engineering for scientists, so the scientists don't have to become software engineering experts. The scientists also need to learn a bit better to ask for the right tools.
The job of a scientist is really not to ship software, that's what a team of engineers would do.
I think that this is the real problem - in academia there is this idea that learning good practices is like a 'dirty' thing that is not required, while instead it would speed up the work and make it more reliable. if you look at chemistry or medicine, there researches have good practices for managing the lab and respect them.
Have you considered that maybe the academics actually know what they are doing?
If you don't write unit tests how do you hedge the possibility of having bugs in your code?
I've taught dozens of grad students enough programming to get the job done and it would have been a total waste of time to make the code that robust. They need experimental results next week, with only one computer ever expected to run the code, not a product demo.
The software isn't their research project, it's a nuisance that they have to deal with. Accordingly they neither want to nor have time to do it perfectly. I cannot blame them.
That said, there should be a system to encourage actual trained programmers to get involved, including coauthorship and consideration in tenure decisions. The current system is bad, I'm just saying it's not the scientists fault here. This is just literally not their domain of interest or expertise, and I would rather they focus on the thing they're uniquely good at.
Their studies / experiments last years.
In CS/ML/Applied Math you sometimes have to write an experiment with a deadline next week. Excuse me if when I'm trying to scramble for a deadline at 3am I don't have my mind toward TDD or I'm not neatly packaging everything in a docker.
> you sometimes have to write an experiment with a deadline next week.
shouldn't happen. And yes, at the moment is like this - sometime you will have to hack. But if the all community start to push for proper practices, instead of just saying "is as it is" - there will be less papers, with more quality.
I think you got me wrong. Shipping quality software is not 'dirty' but requires a specialised focus. One can not do everything by yourself - science and engineering are complementary skills. In your example of chemistry, the chemist who designs a molecule does not spend time to ship the molecule to the world.
I'm sure you don't really mean that premodern societies were more selfless and humble, but that's how you writing can be interpreted. Family name, fateralism and class based society puts even stronger incentives to skew the the truth than individualism.
I do not believe this was any better for the truth in general, as indeed in-group/out-group worldview can be a strong incentive to all sorry of lies. Except in those specific cases where one has to pick between truth that's the best interest of the scientific society and falsehood that's the best interest of the ego.
There is no shame in making mistakes, and being honest about it should take us further as a society.
Questions around methodology abound but at its core this is science walking tall. Not hobbling along loaded down with sugar-coated lies about to collapse into a coma. That this is seen as an exception or extraordinary is quite illuminating in itself.
So, it happens
Maybe with a strong effect for every single subject a little more skepticism would have been warranted in the first place? Some manual spot checking if possible, or using a minimal independent implementation of the analysis code?
Who knows if she'd gotten her grant, her assistant professorship without the publication of this incorrect finding. Who knows who didn't get any of that because they were a bit too careful in their work.
On the other hand, if she hadn't wasted her time on this useless study she might have done more useful studies and her career would be better than it is right now. She might have gotten even better grants if she hadn't made this error, and maybe fewer other people would have gotten grants. Maybe those other people being more careful helped them rather than hurt them.
I don't see how speculating like this is very useful.
I'm saying the downside of making the mistake (for her) might be larger than the upside of making the mistake. I don't think that's obviously wrong.
But my real main point is that speculation of this kind (in either direction) isn't very productive.
It is easy to think over these lines in hindsight[1], but it is much harder to do it when there are no known mistakes. Obviously they had some plausible hypothesis, which predicted and explained results. The more strong result is, the better for hypothesis.
I mean it was possible and maybe wise to suspect bug because results are too good. But it is hard from a cognitive standpoint. She describes bug as hard to find even after she found that results do not reproduce. It was even more hard to find this bug when there were less reasons to believe that there is a bug.
I'm sorry but you seem to have skipped reading the main part of the article. The paper was not retracted:
"The editor and publisher were understanding and ultimately opted not to retract the paper but to instead publish a revised version of the article, linked to from the original paper, with the results section updated to reflect the true (opposite) results."
I know this isn't really what the article is about, but scientists are allowed to be wrong unless and until politics is involved.
[0] https://www.scientificamerican.com/article/italian-scientist...
Like don't be wrong about that. Similar to how you should not be wrong about planes falling out of the sky, bridges collapsing, or medicines killing all the patients.
Kudos to the author for doing the right thing, but the fact that there seems to be no way to remove a paper that is blatantly false because retractions are reserved for deliberate misconduct is horrifying. Not only does this setup long term fucked up incentives (no downside to fraud if you paint it as a whoopsie), but it also harms all work that had cited that work and anyone doing literature reviews not realizing the ground other papers were standing on has dissolved away.
I don't see what the problem is.
Not retracting hid the turd from bibliometric tools that could have easily notifies you of poisoned papers.
The paper authors made a mistake, fine. But the scientific process and peer review process should have caught it. It didn't. The author caught it accidentally and then luckily decided to come forward (bravo!). This begs the question of robustness of the whole scientific publishing process. I hope they adopt the practice of doing a blameless RCA and improve the scientific and academic peer review process.
It raises the question. Begging the question is an unrelated logical fallacy. Unfortunately, there have been a ton of examples of the peer-review process being essentially useless. Things like people deliberately putting things in to test if the paper is even being read and none of the reviewers notice it.
No it isn't, that is a common error in modern English. Raising the question means bringing a question into focus. Begging the question means to assume the conclusion is correct in the premise.
https://www.merriam-webster.com/words-at-play/beg-the-questi...
which seems to indicate my usage is quite acceptable in modern English.
That information is also present in the Wikipedia link.
This is exactly how I'd expect something like this to work -- the author isn't a bad person because they made an error. The co-authors aren't bad people because they failed to catch it. Software and science are hard, mistakes are going to happen.
If anything, I think the researchers learned valuable lessons, and are better researchers as a result. They have an anecdote they can share with more junior researchers about this frightening thing that happened to them, and use that to grow more people.
We should celebrate people who take the time to handle their mistakes properly and share the lessons openly.
But building a year (years?) around a project, attending multiple conferences for it, bringing up a couple students on a research idea, then finding it's all a software bug? That's a nightmare.
That gives you a web page where you can give the full URL to the paywalled one, and it'll give back an "https://archive.xx/address" URL where you can see the full thing.
eg: https://archive.vn/0ikco in this instance
Ideally, that would have been an uh-oh moment.
They did have a null hypothesis and a control group. The problem was that the non-control was different from the control in a way they were careless with, that the scientific methodology does not catch (that peer review and replication ought to have caught). e.g. if the favorable results were better explained by subjects seeing soup taken out of the pot vs. from the can rather than the temperature itself.
It is absolutely possible to find people with an abnormal number of fingers or toes.
https://en.wikipedia.org/wiki/Oligodactyly
https://en.wikipedia.org/wiki/Polydactyly
If 100% of the studied humans have 10 fingers and 10 toes, something's wrong with the study. In this case, the problem is the extremely small sample size: it simply doesn't represent the general human population. In the article author's case, a programming error caused these results.
0% and 100% are always suspect no matter the context they appear in.
I would hand back a grant that was given due to a false result that occurred because of my error. Julia should be praised for doing the right thing with the paper, I just think she should think about going further.
It would be more like losing the remaining payments of a multi-year contract because it gets broken. Or in your analogy, like losing your job.
Sometimes there are going to be mistakes.
But if you sold me software to solve a problem and I paid you for it. You wouldn't refund me if the software didn't solve my original problem? I feel like that's scam but I could be wrong. I didn't get what I was promised and hence I am entitled to a refund.
If you got a laptop with broken keyboard or the privacy friendly app that you paid extra for turned out to be facebook, would you not want your money back?
You can't just use mistake as an excuse if you promise "specific" things like written in your implicit/explicit contract.
Many countries have pro-consumer laws which protect them from exactly this.
That this researcher was honest makes their future research twice as productive in my eyes. Not only are they less likely to make that kind of mistake again, you can probably trust them to not lie about their research.
I think you are correct to want remedy. A big fix, a refund, a workaround, an apology, a million dollars, or a withdrawal of the study. Depending on circumstances.
In my world, most of the implicit contracts are. Shit happens and try to fix it. Yours might be refund me.
Apart from that many great discoveries were made based on initial mistakes.
Actually, encountering this I think I suddenly understand why people get upset about Kickstarter projects not delivering. It is a misunderstanding of a probabilistic situation.
There is a fundamental issue here that is not being discussed is that is the science funding system incentivises people rushing results out and not check if they are real. If Julia had been more careful she would have found the results were false, not published, and not received a grant. So much of the problems with a lack of reproducibility in science is the direct result of people rushing to get out “novel” data and not checking if it is real because if they do they don’t get funded.
Yes, someone whose conclusion was not due to error in programming may have gone without a grant. That's perfectly all right. We optimize policy in the aggregate.
You can make an argument that says that we're not correctly placed on the manifold but no single situation will be convincing for that and you'll need to make some case that a policy change to move to some other point in this space will yield a positive total improvement.
What you are asking for is the equivalent of saying "firing programmers when a bug is found will result in better software"
In science it will result in much less interesting science being done, because this will desincentivize risk.
The big problem in science funding is that it already incentivizes low risk projects (you should have preliminary results that proof it works...) so let's not make it worse
From a software perspective it is like paying developers by the number of lines of code delivered and telling them to not worry about bugs. Not an approach that is likely to deliver quality software.
What we need in science is to find a way that people have the time and incentives to publish results that are as accurate as possible.
Somewhat off topic I do wonder if we would get better software if we fired programmers if bugs were found. I suspect we would not get much code written, but the code that was written would be very high quality.