A Plan Made to Shield Big Tobacco from Facts Is Now EPA Policy
nytimes.com
nytimes.com
> The new rule, public health experts and medical organizations said, essentially blocks the use of population studies in which subjects offer medical histories, lifestyle information and other personal data only on the condition of privacy. Such studies have served as the scientific underpinnings of some of the most important clean air and water regulations of the past half century.
because I'm leaning (b) on principle from my own analysis of the tradeoff but there aren't powerful sarcasm flags in this post that I see, nor is there much in the way of arguments to convince me
The typical counterargument is that insurance providers would maliciously charge higher based on what they see. However, that’s silly since they already have that data today.
What would showing off private data do? Put a minimum age on smoking? Stop it from being advertised to minors? Ban ads in a majority of media? Put health warnings on packages? Put high taxes on cigarettes?
Oh wait, we already have that.
Short of prohibition, what new law do you want at the price of health data privacy? Health data is pretty much our last legally appointed sacred data point in the US and it's already weak enough as is. HIPAA is great and all, but it's still not enforced that great. At that since we are decriminalizing weed and alcohol prohibition didnt work... no. Everyone's health data is not a fucking fair price to pay for your short sighted emotional knee jerk.
Something like regulations to reduce lead in the air could be the biggest material improvement in people's lives since we eliminated lead from paint and gas. Bigger than any improvements we've made in education directly. But rules like these just get in the way.
Open data has provided tremendous value in recent years, not just for transparency and civic engagement, but for innovation and progress. The fact that we can find large, detailed datasets for topics like sports and birdwatching but not for health is a travesty.
1. Make the underlying model public (e.g., you ran such and such regression, here’s what you included and here’s a table of coefficients). Inarguably a good idea. Totally unobjectionable. You should correctly be suspicious if the model (and the R, Stata, SAS or whatever code to produce it) is not available.
2. Make the underlying data public. Completely different! I have done research with (e.g.) detailed birth certificate level data. Such data are confidential by state and federal laws generally. Other researchers work with detailed income histories from the IRS. As another example, I have worked with electronic health records, which include names, addresses, health conditions, etc. Almost no one thinks this stuff should be public.
It is fair IMO to require researchers to describe what data they use and how someone else might get it (even if that includes the line “find a collaborator at the IRS”). But it is unreasonable to require researchers to make data itself public, when they may not have the right to do so.
What’s worse is that the reason this regulation is being proposed is to prevent research informed by private, confidential health data (eg birth certs) from being used to make public policy, especially in areas around pollution and environmental regulation.
Concretely there's applied ML papers on confidential military data, where I'm fairly sure the results are just bogus or cherry-picked or done with incorrect procedure, but with the data being unobtainable (even to those with clearance the data isn't materialized in some DoD database for reproducing) it makes the published results entirely unreproducible and unchallengeable.
People already do this when there's basically nothing at stake though.
Policy has to be based on something. Public research is being attacked because it's not perfect. Okay. If you're going to do that, you should apply the same expectations to private research, and also require that firms prove that their products are safe, to the same rigorous standards that you expect out of public researchers.
I don't see how that's any different than when congressmen get interviewed by the talking heads and say "the research is clear, there's been study X by Y and P by Q that show that what we need to do is enact <insert some asinine extremist legislation that that congressmen supports>". They're both lying in a plausibly deniable manner in order to market something to the public. Anyone who bothers to verify either claim will find it's all crap.
The problem isn't the research. It's that we blindly trust "the research" and people know this so the research gets manipulated and cherry picked and whatnot. Two hundred years ago these sorts of people and entities all invoked gods name in order to peddle crap. I have no idea what the solution is but I think whatever it is will involve the public being more disapproving of parties that behave in a slimy manner like this.
One of the two parties is currently seriously considering not accepting the results of an election they lost. The leader of that party is busy calling up election officials, and threatening them into magically finding thousands of votes that will give him the victory.
And the rank and file of the party isn't distancing themselves from this behavior. Instead, they are turning their ire at the person who leaked it.
On the order of slimy political behavior that voters don't punish politicians for, this issue we're discussing won't even register. We're currently so far down the rabbit hole, I don't think we'll ever find our way out.
> I have no idea what the solution is
That's because there is no perfect solution. We don't live in a perfect world. We just have to find the best solution out of a space of sub-optimal ones. And in this case, public research is a much better starting point for policy than private research, or no research (which is what would happen if we applied your standard.)
You have two sets of incomplete data. One from public research, where some of the data is missing, and one from private research, where nearly all of the data is missing.
Which do you base policy on? Doing nothing, by the way, is also making a policy decision.
I will also point out that changing the rules post-facto, like we're doing here, is like asking an open source project to verify that every single contributor to it has consented that the project migrate to a new, incompatible license.
They likely assume that any regulation will be bad by default unless it's proven to them -- and won't put any effort into reading that proof when offered.
Software engineers get used to the notion that things either work or they don't. Policy is a much messier process, where you do the best that you can with incomplete information. And software engineers are generally comfortable enough under the status quo that they will set high standards for change -- and aren't interested in how that affects anybody not in their immediate circle.
Did you mean NY Times?
I too also hate when this sort of language is used. Typically it's "Twitter thinks" or "Facebook reacts" when the correct terminology should be "These 3 users on Twitter" or "This specific person on Facebook".
80k views is nothing to sniff at, but that's hardly "viral".
Not really the same as what you're talking about, but I was reminded because it's another example of lazy use of language by journalists.
I think we're in agreement here; stop lazy journalism.