No funding for uncomfortable results
johndcook.com
johndcook.com
Anyway, Gillberg the professor who led the study had promised all participating families total anonymity. But in the mid nineties, other researchers started to doubt the veracity of Gillberg's results. They thought they had found flaws in his methodology and wanted access to the raw material to double-check his results. Gillberg refused, realizing that anonymizing the data sufficiently so as to not leak any identifying details would be impossible.
Eventually, the other researchers went to court and got a court order demanding that Gillberg give up his data. At which point he destroyed all the raw data thereby completely removing any scientific underpinnings the DAMP diagnosis had. There is a lot more to be said about this controversy, but one thing is clear and that is that you shouldn't promise anonymity if you can't keep it.
Although it sounds like he did keep it, at significant cost.
“In the few experiments a decade ago where publication would have been possible, academic journals refused to do so for reasons having nothing to do with the scientific quality of the work. Computer-science publications refused to publish re-identification experiments unless the paper also included a technological solution, notwithstanding assertions that publishing these experiments would inspire technological innovation to address the real-world problem. Health-policy publications refused to publish re-identification experiments related to health data from fear that reaction might make data sharing more difficult, despite assertions that because technology was fostering unprecedented levels of data-sharing, it was timely to scientifically re-examine data-sharing practices. Even my Weld example and related demographic analyses, despite making significant contributions to privacy regulations worldwide, were refused publication by more than 20 academic publications at the time.”
Personally if feels like the work was an important journalistic effort, but its scientific merit is less clear.
But looking at what was done, it doesn’t seem novel. It’s just combining two datasets, and while journalistically important, doesn’t seem to be sufficiently novel for publication in most journals.
The question is why you think novelty is relevant. Replication studies aren't novel but are critically important to real science. Novelty is not and should not be mandatory.
For a healthcare replication study, I would guess you be hard pressed to get a replication study published for the same demographic, and same sample size as a previous study.
Yes, but this shouldn't be the case. The replication crisis clearly points out the problems with this.
This is exactly right - Dr Sweeney went on to cofound a PhD program at CMU focused on this intersection, called "Computation, Organizations and Society". It's not just that policy folks aren't tech literate and that applied cs folks don't know law -- there are interesting technical results that evince at the intersection, like differential privacy frameworks and mechanism design.
A challenge of this kind of work is that it's difficult to prove generality or prediction power, because you are studying distributed systems that don't repeat themselves. Economists and sociologists are used to this problem, but it's hard to intersect with the more technical applied sciences that you would like to vet the work. So most practitioners have to pick one side or the other (hard vs "soft" communities) one publication at a time.
You can't solve a problem without discovering what the problem is.
> Over 20 journals turned down her paper on the Weld study ...
And this:
> A decade ago, funding sources refused to fund re-identification experiments unless there was a promise that results would likely show that no risk existed or that all problems could be solved by some promising new theoretical technology under development.
And she didn't just do a "back-of-the-envelope calculation". She re-identified an actual public dataset, including the then governor of Massachusetts :)
I’d like to see the rejections, my guess is they mostly said “this is not novel or unexpected”.
To make the work novel you might suggest a new approach to anonymizing data for example...
Well, given how widely vulnerable de-identification was being implemented, that seems unlikely.
And yes, showing this with a real dataset is nice, since it validates the estimation. But it still doesn't explain why it's worthy to fund. Funding is very competitive, and generally goes towards prospective research that can lead to innovation in the field. The article suggests that funding should go towards already-finished descriptive dataset analytics confirming back-of-envelope estimations, with no clear plan to find a solution.
It's also possible that this article is not expressing the nuance behind the actual work and reasons behind the rejections, so I'm not judging the researcher's work, but rather judging the claims in this article.
Though even so, your comment still stands (since I'm assuming whomever is in an editorial position has some kind of science background, at least at some of those journals).
As a social scientist, I'm not so sure...
It reminds me of the times when people who aren't programmers seem to make this discovery that there is such thing as technical debt, and that programmers should be focused on programming in a way that avoids technical debt ("these programmers should focus on writing maintainable code! Not just short term hacks"). It's an oversimplication of the situation that is not helpful.