EPA's proposed new rule would allow it to ignore the best available science
blogs.scientificamerican.com
blogs.scientificamerican.com
https://www.federalregister.gov/documents/2018/04/30/2018-09...
I think it sounds great. Reproducability and verifiability are extremely import in scientific research. Unlike what this article and some people here claim, it does make provisions for protecting the privacy of users. I also don't see anything hinting that it will affect past studies retroactively.
My only complaint is that it targets the EPA specifically, which is pretty suspicious with the current political situation. It would be great if this affected all studies which receive public funding.
I don't mean to sound anti-corprate, but environmental regulations only exist because companies externalize costs onto society and this increases the burden for balancing exposing that cost.
If so I don’t disagree.
I disagree with this article stating that raw data needs to be kept secret. It’s pure bullshit and anti-science.
I cannot tell from your tone and the jab at “technical geniuses” if you’re saying this is a bad thing.
Yes, environmental protection costs money. This money should be paid by the companies who are directly or potentially placing the environment at risk. This burden should not be externalized onto the public so private entities can make more money.
As far as I can see that's the only immediate information this article bothered to speak about the rule itself. Anybody see a summary on this rule to share with the crowd? Or did I miss something?
EPA is proposing to eliminate consideration of scientific studies unless the raw study data are made publicly available — possibly including participants’ personal, confidential, and private information — data that may be subject to legal, ethical, and human subject research protections.
This requirement would eliminate human health studies that use integral medical, lifestyle, and geographic data, as well as studies that include confidential business information, such as information from studies conducted by industry to demonstrate safety of a pesticide or other toxic chemical. It could also eliminate consideration of older studies for which data are no longer available or accessible, even if the data have been reanalyzed and the studies have been validated, replicated, reproduced, and undergone rigorous and independent peer review.
Far from using the best available science for EPA decision-making, the proposal would severely limit the studies and scientific evidence the agency would use to fulfill is statutory mission to protect human health and the environment. This is not a new idea. It is the result of a decades-long campaign to undercut reliance on groundbreaking research...
The proposed regulation provides that when EPA develops regulations, including regulations for which the public is likely to bear the cost of compliance, with regard to those scientific studies that are pivotal to the action being taken, EPA should ensure that the data underlying those are publicly available in a manner sufficient for independent validation.[1]
Seem eminently reasonable to me. And seems aligned with the "open science" movement to me.
And in terms of privacy...
EPA believes that concerns about access to confidential or private information can, in many cases, be addressed through the application of solutions commonly in use across some parts of the Federal government. Nothing in the proposed rule compels the disclosure of any confidential or private information in a manner that violates applicable legal and ethical protections.
[1] https://www.federalregister.gov/documents/2018/04/30/2018-09...
In terms of new research are you saying that it’s ok to keep the raw data a secret? Why? We don’t do that for any other type of science.
Yes, the best way to prove reproducibility is to reproduce it, but that can’t always be done and at any rate is no excuse for keeping raw data a secret.
This is blatant disingenuousness. What they are saying is "you need not disclose private information, but we won't use your study if you don't".
Also, please don't be so credulous. This rule has been promulgated, invented, and proagandized by people who are trying to hamstring the EPA from using scientific studies to regulate industry. It's not new, it's just newly empowered.
You appear to be falling for the plausible cover story, which is disinformation. People are really bad at being able to discern the intent behind "plausibly good" organizations and movements.
Suppose zero studies that satisfy that criteria study lead. Should the EPA suddenly say lead must be safe in any concentration?
Good luck tracking down data from ~50 year old studies published in 1970. Further, their is little point in replicating old studies so often that's all we got.
Try and publish a paper in Nature and tell them “I can’t share the data”.
It's not a bad goal, though making data public is always a bit tricky when data collected from human beings is involved.
>With this notice, EPA is soliciting public comment on a proposed regulation designed to provide a mechanism to increase access to dose response data and models underlying pivotal regulatory science in a manner consistent with statutory requirements for protection of privacy and confidentiality of research participants, protection of proprietary data and confidential business information, and other compelling interests.
>The proposal takes comment on how to ensure that, over time, more of the data and models underlying the science that informs regulatory decisions (over and above the dose response data and models underlying “pivotal regulatory science”) is available to the public for validation [13] in a manner that honors legal and ethical obligations to reduce the risks of unauthorized disclosure and re-identification.
>EPA believes that concerns about access to confidential or private information can, in many cases, be addressed through the application of solutions commonly in use across some parts of the Federal government.[16] Nothing in the proposed rule compels Start Printed Page 18771the disclosure of any confidential or private information in a manner that violates applicable legal and ethical protections. Other federal agencies have developed tools and methods to de-identify private information for a variety of disciplines.[17] The National Academies have noted that simple data masking, coding, and de-identification techniques have been developed over the last half century and that “Nothing in the past suggests that increasing access to research data without damage to privacy and confidentiality rights is beyond scientific reach.” [18] More recently, both the National Academies and the Bipartisan Commission on Evidence Based Policy [19] have discussed the challenges and opportunities for facilitating to secure access to confidential data for non-government analysts.
>Considering the breadth of dose response data and models used in the development of significant EPA regulations, the requirements for availability may differ. These mechanisms may range from deposition in public data repositories, consistent with requirements for many scientific journals,[20] to, for certain types of information, controlled access in federal research data centers that facilitate secondary research use by the public.[21] EPA should collaborate with other federal agencies to identify strategies to protect confidential and private information in any circumstance in which it is making information publicly available. These strategies should be cost-effective and may also include: Requiring applications for access; restricting access to data for the purposes of replication, validation, and sensitivity evaluation; establishing physical controls on data storage; online training for researchers; and nondisclosure agreements.[22]
https://www.federalregister.gov/documents/2018/04/30/2018-09...
Hope that helps.
The proposed regulation provides that, for the science pivotal to its significant regulatory actions, EPA will ensure that the data and models underlying the science is publicly available
The public availability requirement sounds nice on paper, but in practice, it means ignoring a lot of science during rule-making because the original data can't be published due to genuine privacy concerns.
There are some common-sense workarounds. For example, requiring an EPA or third-party audit of the original dataset and analysis. Or simply carving out exceptions for research whenever it can be demonstrated that an IRB (and therefore federal law) would not have allowed the research without restricting public access to data/analysis.
It is more limited if you read the whole sentence, which goes on to qualify which data they are interested in (specifically "publicly available in a manner sufficient for validation and analysis.").
These rules are a hot political topic that are years in the making. The people pushing them are not scientists and have never expressed genuine interest in reproducibility outside of this one issue. And scientists who have spent years fighting for reproducibility have publicly stated they oppose this rule. And all the major proponents have strong anti-regulatory preferences.
These rules are not about reproducibility. This rule is an intentionally crafted Catch-22 designed to exclude medical research from consideration so it's easier to justify repealing or not implementing environmental regulations. Otherwise, the rules would find a middle ground; e.g., by allowing the EPA access to relevant scientific data and analysis without risking https://en.wikipedia.org/wiki/De-anonymization by posting sensitive medical information about thousands of patients on the public internet.
> Otherwise, the rules would find a middle ground;
They actually did, if you read the actual proposal, which talks about disclosure in ways that enough for validation but does not infringe privacy. But you seem to have arrived to pre-determined conclusion already, and there in fact can be no middle ground that would be acceptable to you, as it seems.
Then you're one of today's lucky 10,000: "Robust De-anonymization of Large Sparse Datasets" http://www.cs.utexas.edu/~shmat/shmat_oak08netflix.pdf
A lot of environmental medical research results in dense and small data instead of sparse and large data, which obviously means it's even easier to de-anonymize.
And the really important take away from the above paper is that you need to know about all the other public data in the universe in order to really ensure anonymity. Even when anonymity seems like it's "obviously" not going to be a problem.
You never know what dataset will be leaked/published tomorrow correlating columns D-F in the spreadsheet you gave EPA to someone's first and last name. So privacy is like computer security: you can never be totally sure. So when the stakes are high you should be extra cautious. And the stakes when sharing medical research are very often high.
> If that's true, how people validate any medical research at all?.. I have very hard time believing that's the only way to do modern medical research known to our civilization.
This isn't exactly a new problem, so scientists have had a little while to figure out how to make peer review work.
You share the data. You share with trusted, relevant third parties who promise to preserve the anonymity of study participants. Just becuase you do not publish possibly de-anonymizable and definitely sensitive medical information on a public website does not mean you don't share the data with interested and relevant third parties.
> But you seem to have arrived to pre-determined conclusion already, and there in fact can be no middle ground that would be acceptable to you, as it seems.
On the contrary, I do literally exactly that in my very first post in this thread: There are some common-sense workarounds. For example, requiring an EPA or third-party audit of the original dataset and analysis. Or simply carving out exceptions for research whenever it can be demonstrated that an IRB (and therefore federal law) would not have allowed the research without restricting public access to data/analysis.
The situation where IRB says "no" to public disclosure should explicitly trigger exemption, with the duty falling back to the agency to prove the IRB was wrong in the first place. The fact that this exemption is not automatic is what makes this a catch-22 that pins researchers in-between their IRB and the EPA administrator in charge of evaluating their exception.
> important take away from the above paper is that you need to know about all the other public data in the universe in order to really ensure anonymity.
I do not see how it's takeway from that paper.
> You never know what dataset will be leaked/published tomorrow correlating columns D-F in the spreadsheet you gave EPA to someone's first and last name.
You also never know whether somebody doesn't just break into your system, downloads all secret data and publishes them. Yes, you have no guarantee against bad actors and mistakes. But if that's the border condition then no research should be performed at all - after all, once the data exists, there's no guarantee it won't somehow get leaked out.
> The situation where IRB says "no" to public disclosure should explicitly trigger exemption
IRB is a part of the institution, right? So you essentially delegating the decision about the exception to (part of) the institution. What the point of having the rule then if every organization can make their own exceptions on demand? And that makes trivially easy to hide the data in sloppy research - just sprinkle some private data on it, show it to IRB, they cry out "no, this can't be released, no way!" - and voila, you are safe from review.
1. Sometimes it is.
2. I already gave substantive reasons why this will commonly occur in environmental medical research -- dense datasets about small populations that capture sensitive information about individuals.
> I do not see how it's takeway from that paper.
Because that's literally the only assumption in the adversarial model other than "access to the published dataset". And correlating with an auxiliary dataset is literally how all actual instantiations of this attack work.
> just sprinkle some private data on it, show it to IRB, they cry out "no, this can't be released, no way!" - and voila, you are safe from review.
1. My proposal doesn't say "IRB happened to say you can't publicly publish this one particular protocol." My proposal says that IRB explicitly judges that no such protocol exists.
2. IRBs are not that arbitrary and are themselves audited.
3. If you're this paranoid about intent, then you can't trust the research anyways. If the research is committing intentional fraud, they could just do it at the data collection step.
The way I see it, either:
A. Even if we can't trust individual scientists, we can more-or-less trust a broad subset of scientists with different and well-aligned motives (as in, a whole group of: the IRB board, the PI, the people reviewing the IRB board) when they say that data can't be shared for privacy reasons; or,
B. We should expect that literally dozens of scientists, all with different motives, will routinely and actively collude to commit wide-spread fraud.
If (A), my approach makes sense. If we max out paranoia and go with (B), then I don't know why we should trust the data regardless of whether it's published publicly...
Considering one of the tenants of the scientific method is reproducibility, how can you reproduce something if the raw data is unavailable?
If you get cancer because of a chemical spill do you want all of your personal information released to help prove that the chemicals spilled poisoned you? No. You want your personal data protected.
Sounds like a red herring to me.
It also mandates reproducibility. It isn't exactly easy or even possible to reproduce a 70-year long study or a study of victims of a disaster which is the point of the American College of Obstetricians and Gynecologists and The American Academy of Pediatrics.
https://www.regulations.gov/document?D=EPA-HQ-OA-2018-0259-2...
Reproducibility is a key part of the scientific method. I'm surprised people are so willing to toss it aside.
I probably sound like a broken record, but this is standard across scientific disciplines. No raw data and people call bullshit.
https://www.vox.com/science-and-health/2017/3/10/14871696/sc...
Scott Pruitt denies basic climate science. But most of the outrage is missing the point. ... It’s not about Pruitt and it’s not about facts. (from 2017)
[Complaints] are all about what Pruitt believes. And in the end, who cares what he believes? He is a functionary, chosen in part to dismantle EPA regulations on greenhouse gases. If it weren’t him, it would be some other functionary.
The goal is to block or reverse any policy that would negatively affect [conservative] donors and supporters, who are drawn disproportionately from carbon-intensive industries and regions. That is the North Star — to protect those constituencies. That means, effectively, blocking any efficacious climate policy (which, almost by definition, will diminish fossil fuels)
No.
>It feels like we’re well on our way into a new dark age.
Bill Gates would probably disagree with you on that (http://bit.ly/2mg08wL)
The groundwork was laid in the US starting in the 1970s. With the founding of Cato, Searle, Olin fnds. Murdoch bought the NY Post in 1976. Fairness doctrine was repealed in 1987 allowing Rush Limbaugh to go bigtime. Fox started in 1997.
I am actually a conservative on many issues, including UBI, work rules, tariffs, caution when making new law and regulations, and appropriate pricing of risk. However the real conservative movement has been redirected as a way to link wealthy donors’ interests with identity politics, for votes. Part and parcel of that is a huge propaganda machine which spouts lies. That propaganda is corrupting the West and I am afraid for its effect on democracy and average people.
> what does that have to do with this regulatory proposal?
Article's subtitle is: The agency’s proposed new rule would allow it to ignore the best available science
There are thousands of old studies for which the data is unavailable, incomplete, was lost, etc. Are you going to pay for all of them to be reproduced? Are you ready to give up penicillin? Surgeons washing their hands between operations? The links between tobacco and cancer? Every single medical advancement (Because patient confidentiality precludes the release of raw data that goes into medical studies)?
I'm all for reproducibility, but this has enormous consequences.
The scenario is not so hypothetical. The circumstances and particulars change, but this situation describes a lot of medical and environmental research.
Should the EPA take the researcher's word? Maybe not. But there's a lot of workable middle ground in-between "blind trust" and "anyone with access to a library computer and a stats textbook can figure out participant 494 -- who math tells us can only possibly be Johnny D. of Smallsville, VA -- has testicular cancer".
And the situation is even more difficult than this because de-anonymization is an active area of research with some surprisingly powerful results published in recent years; i.e., it is highly non-trivial to determine if a dataset is de-anonymizable. And the answer always depends on what data about a person's identity is publicly available, which especially in the age of annual massive data leaks change from year to year. So even if researchers don't think it's possible to de-anonymize study participants using today's mathematical tools and today's public datasets, they'll often choose to keep the dataset private just in case.
Have the authors of this article actually confirmed that this would include non-anonymized data unfit for public consumption?
Added: The actual document does not specify "raw data", it is much more specific about what data they expect.
> The proposed regulation provides that, for the science pivotal to its significant regulatory actions, EPA will ensure that the data and models underlying the science is publicly available in a manner sufficient for validation and analysis.
"A manner sufficient for validation and analysis" is a completely different standard than "all the data the study produced".