167 karma · joined March 14, 2014
You said that Lesné was "not that highly cited". But his main fraudulent paper was cited 2,300 times, making it the fifth most highly cited Alzheimer's paper since 2006! [1]
Berislav Zlokovic's likely fraudulent papers were cited 11,500 times! [2]
It's hard to imagine these highly papers didn't redirect at least some scientists to do pointless followup studies. Of course, in the counterfactual world the scientists might still have been doing pointless studies, but we'll never know...
[1] https://www.science.org/content/article/potential-fabricatio... [2] https://www.science.org/content/article/misconduct-concerns-...
I wrote a blog post on how to make this easier, including a new criminal statute specifically tailored for scientific fraud. https://news.ycombinator.com/item?id=41672599
How much time could have been saved towards an effective treatment? It could be as high as a decade, but of course more likely it was zero years. I averaged it out to 1 year.
Now suppose you think that 1 year is orders of magnitude too high, and that in expectation it averages out to a 1 day delay. Even then, I estimate 100,000 QALYs would be lost, making this a tragically high impact case of misconduct.
--- Final point: Nobody doubts that science is error correcting. The point is that the errors are corrected far too slowly and many never get corrected at all. It's incredibly hard to develop good theories when you know that 30-50% of the results in your lit review are false.
From the article:
> One-third of major storms arrive unexpectedly, according to the SWPC’s own 2010 analysis. And that’s not just the small storms. According to a news article in Science, the SWPC might be also be poor at identifying the characteristics of severe storms, since they are so rare.
By the way, dplyr doesn't use multi-indexes. I actually think this one of the reasons (although not the biggest reason) dplyr is easier to use.
http://pandas.pydata.org/pandas-docs/version/0.18.0/whatsnew...
There have been a few independent attempts to add dplyr-like functionality to pandas without being backwards incompatible (e.g. dplython). I'd be very happy if the core pandas team went down this path.
That being said, I don't have a good understanding of how strong the distinction is between "design issues" and "issues where money helps". There must be some overlap.
https://twitter.com/Chris_Said/status/715249097326768128
https://twitter.com/Chris_Said/status/861244045535756290
I could be wrong, but I'm pretty sure that these would be solved by pandas API design improvements, not with numpy improvements under the hood. (NB: As always, a big thanks to the developers for all their work.)
That being said, I do wonder if numpy is the most appropriate recipient. In my experience with data science, the tool that would benefit the most is not numpy, but pandas. While data scientists rarely use numpy directly, every data scientist I know who uses pandas says they are constantly having to google how to do things due to a somewhat confusing and inconsistent API. I use pandas at work every day and I'm always looking stuff up, particularly when it comes to confusing multi-indexes. In contrast, I rarely use R's dplyr at work, but the API is so natural that I hardly ever need to look things up. I would love if pandas could make a full-throated commitment to a more dplyr-like API.
Nothing against pandas -- I know the devs are selflessly working very hard hard. It's just that it seems there is more bang for the buck there.
It would be interesting though to have people try to guess suitable values of m and C and then see how close their MSEs get to the James-Stein MSE. I suspect that some people's guesses would be meaningfully off target.
Regarding MCMC, one of the things I try to emphasize throughout the post is that the best solution depends on your needs (for example if you want a full posterior). In fact, most of the post is devoted to quick and simple methods -- not MCMC -- because they are good enough for most purposes. I welcome your feedback though on how I could make this point clearer.
TL;DR: One blog post is for Rotten Tomatoes and the other is for Metacritic.