Covid-19 vaccines and treatments: we must have raw data, now
bmj.com
bmj.com
“Raw data” is submitted to the FDA in the CDISC format. This format contains a lot of pretty sensitive medical information, including which diseases a patient has, their medical history, what drugs they take, etc. This is supposed to be anonymized, but if the public were to be able to get this info, I strongly believe there is enough information to re-identify patients. And it’s not as simple as just removing the sensitive medical data because the primary or secondary analyses may be dependent on them.
https://www.hhs.gov/hipaa/for-professionals/privacy/special-...
It's also possible to reach an outcome that the data can be shared but must be de-identified in a way that precludes important statistical tests. Any sparse feature is incredibly powerful for re-id, and combined with other features might be difficult to share without running afoul of de-id best practice. The problem: rare medical conditions that must be accounted for in a statistical study of vaccine side-effects are examples of sparse features. So you can share the dataset, but not in a way that's useful for a non-GIGO statistical study.
You also keep ignoring the issue of consent. Step one is to ask patients.
Or the person shared their story on the public website of a "Run For The Cure" style website about that genetic disorder.
Or so on.
"Sparse Features" aren't always a thing that the person wants to keep private.
https://vaers.hhs.gov/data/datasets.html
This is the type of raw records we are talking about, and they carry useful information.
Yes, maybe the dates could be anonymised better, if I know the age of someone, and I know when he got his vaccine precisely, and when he got his MRI and why he did it, then maybe I can find and connect back the record. But how likely this attack can be done in practice, and why would you do it to someone who already disclosed part of their records, plus how does it scale ?
However, the benefits are really present, it's not sharing just for the sake of sharing but to bring advance to the global research and knowledge of medicine.
In practice, who is it ? I don't know, who can know ? Doctors who can tie it to other pieces of information that we don't have...
However, in the meantime, the researchers and doctors who don't have these foreign keys, they can totally do interesting and useful discoveries about drug interactions and vaccine side-effects.
Even if you take more care than AOL did in anonymizing your data, the unfortunate reality is that any publication of data increases the knowledge an adversary has at identifying somebody. Anonymizing is more about reducing the chance someone is identified than guaranteeing they never will be. And high dimensional data is particularly hard to do so in a way that retains the data's usefulness.
33 bits is an old defunct blog on this topic, but it has some interesting posts and academic papers if you want to go down the rabbit hole: https://33bits.wordpress.com/about/
Specific paper on Netflix deanonymization: https://33bits.wordpress.com/about/netflix-paper-home-page/
This paper is a very good starting point:
There are ways to inject randomness into a dataset, giving subjects plausible deniability without compromising the integrity of the outputs. See https://pair.withgoogle.com/explorables/anonymization/ for a nice example.
Bottom line is, there are ways to share such data safely. Whether the data owner cares enough to do the extra work, especially if doing so removes a competitive advantage, is another story.
Furthermore, the FDA submission dataset is essentially a database with dozens of tables each with often hundreds of columns. It’s a LOT of data, with exponential complexity to make sure all the right fields are redacted. There’s also the point that pharma companies are under no obligation to release this data. It’s generally considered proprietary. That said, due to the substantial amount of government funding provided to the development of the vaccine, I think we should be entitled to this information.
Nobody's saying "tamper with the evidence to demonstrate favourable outcomes" here.
And if tampering with the data to favour the data donor were a concern in the first place, then it should be a concern regardless of whether the data donor said they followed further anonymization protocols or not.
Yeah, so you have badly misread my point, and you're arguing against something completely unrelated to what I'm saying.
Of course you can add randomness to the results to hide individuals without messing with the results in aggregate. That's obvious, and I'm not disputing it. The issue is that that's a PR nightmare. Nobody here might be saying "tamper with it to change the outcome", but I guarantee you that people out wider media will accuse the CDC and Pfizer of doing exactly that. I can see the chyron now, "CDC ADMITS TO MANIPULATING DATA".
Organizations like the CDC need to think not just about how their message will be heard, but how it might be twisted and abused by bad faith actors. You are absolutely right about how you can safely introduce randomness, but you are completely wrong about how that will be handled by the public after bad faith actors get their hands on that fact.
But I don't necessarily think this would go down much worse than "Big Pharma refuses to share data, what could they be hiding from us" bad faith actors.
At least in the 'inject randomness' scenario you can do some damage control in the media by bringing in experts to explain how this is a valid thing to do and how it protects the average joe from having their lives turned upside down just because they foolishly agreed (or worse, did not agree) to have their data collected.
I get that some people want to prevent the use of the data for discrimination...just legislate that then.
At the time when you inject hundreds of millions of people with it, I'd say this should be mandatory to disclose as much information as possible regarding the actual trials. I don't trust the FDA one bit.
Where is the bad thing in you posting this information here right now?
Actually, that's not even analogous. A better analogy would be you post all this info under a pseudonym -- e.g., "rvnx" or something like that -- and then the mods doxing you without your consent.
For example, having the list of symptoms, medications and even (just an imaginary example) pulmonary X-Ray or also on the chest having an open dataset of MRI pictures with the age of the patient, history of diseases would considerably help and there is no way you can link it back unless you know the history of the patient already (but then, what's the issue; none because you already know the information).
Plus if you ask the person about it, I really don't see the problem :|
Yes in theory, there is a possibility that someone has a database somewhere of all the boobs spacing, and can determine approximate lungs size, try to guess what could be the likely owner of a lung picture, etc etc, all of that for what ?
The person who is going to benefit from it, is clearly the COVID-19 researchers, and the end-patient.
If you were to have access to my MRIs you could almost certainly identify me based on physical characteristics which can be seen in my MRI data which would correlate with data found on social media where I've sometimes generally discussed my medical conditions without going into specifics. There are MRIs for conditions I haven't discussed, in addition to the identifying ones I have.
There are a lot of people with fairly unique medical conditions and body situations (eg: missing a limb, surgical implants). The gap between "I got hurt in this way" and "here's the MRI scans of my body in detail along with other lab notes and other unrelated and possibly extremely embarrassing conditions" is huge
I have deposited datasets with full transcriptome data, which you can use to infer most of the genome in a straightforward manner, from clinical trials. In that case, you simply use off-the-shelf tools to anonymize datasets by flipping lots of SNPs. Additionally, raw data is made accessible only through an ethics committee. Summary counts from transcriptomes are publicly available, though. There are equivalent practices for other biomarkers.
I must say The British Biomedical Journal (BMJ) is really spearheading all transparency efforts. They have been publishing lots of inconvenient truths for many years. For example, they reported some fairly serious conflicts of interest where professors from Oxford are getting huge "consulting" fees from pharma. And they pretty much publish any reasonable letter to the editor. Nature, Science and Cell, on the other hand, are heavily biased.
Right.
1. I'm not sure you understand the extreme position being taken by rvnx in this thread.
2. Presumably the patients agreed that their transcriptome data could be shared in this way, and you didn't change your data handling post-facto without their consent.
Said agreements have previously gone through ethics committees which tend to be extremely strict. They are rarely changed afterwards, since this requires agreement of patients, which is impractical to obtain in all but tiny studies.
The only exception I can think of is Scandinavia. They have national registries where all census information, plus electronic health records, is available to researchers by default, unless you explicitly opt out. Their philosophy is that this is a negligible reduction in privacy, in exchange for medical progress. I must note said records are not publicly open, you need to apply for access.
As a consequence, Sweden and Denmark have some of the best epidemiological research in the world and they've been able to spot pretty odd correlations (e.g. ibuprofen and male infertility).
I have had helped draft at least 6 studies with materials coming from our own clinical trials and deposited all data myself. Went through institution review boards without an issue, just like many other equivalent studies.
Perhaps your experience is different because of location? I'm referring to EU.
AFAIK this is also possible in the US. Probably more hoops to jump through? I don't know anything about the EU. (But also this discussion is basically totally unrelated to the original article or this thread.)
Your comment is somehow relevant to the article and insightful wrt the topic but also off-topic in this thread.
The type data you are suggesting should be shared -- and, more importantly, the consent assumptions -- are not related to the current discussion.
Until/unless patients agree otherwise.
I.e., the status quo.
Otherwise it's sort of like my 10-year-old wanting to help me replace an engine in a car. It's important to let him do it so he can learn but it's likely going to take me much longer to on-board him and get him up to speed than if I did it myself.
No. Intent doesn't matter. That's the whole point.
To wit: there's a reason that reidentification attacks are called attacks! No one intends to be hacked, and writing a program that you intend to be secure doesn't imply that the program is in fact secure. So too with anonymization.
> having an open dataset of MRI pictures
But we're not talking about MRI pictures, are we?
OP strongly believes that, in this case, there is enough information to re-identify patients. I agree.
Plan A: - sit and do nothing "it's impossible to release any data, too bad, we'll never know!"
Plan B: - "we have a risk that a few % of ppl may be identified if we already know their medical history/details, it's impossible to release any data, too bad!"
Plan C: - "we remove the biggest identifiers to make sure the majority of the people cannot be reasonably identified, this is going to take a few days or weeks of work but will bring benefits to the overall population and transparency on the treatments"
I'm fully supportive of the right to pseudonymous speech, but it strikes me that you are in an awfully awkward position of speaking anonymously while demanding that you should have access a trove of patient history without signing any sort of data use agreement and without even the most basic step of getting permission from the people whose patient history you insist on seeing.
Asking the user:
"This record is going to be shared publicly, do you agree to share this information (without revealing your name and address) ?
> [...]
This information can help researchers, doctors and people interested into science to discover new information about drugs and diseases interaction."
We learnt that somewhere there is a man/woman that is depressed and has weight management issue.
We can take policy actions in order to improve the situation and monitor how this affects life of people (e.g. how much antidepressants are used in a year, are the food policies effective, what conditions are often linked to sudden cardiac arrest, etc).
And "Rebecca", nobody knows it's her, except herself, and those to whom she decide to publicly disclose her conditions (e.g. on social networks or to specialists).
Asking people in the study if it’s OK to release their data means the resulting data set is skewed. People with certain conditions and other variables will be more or less likely to agree, leaving a non-representative sample of the actual study.
It will no longer be random.
Actually, set aside consent entirely.
The problem is NOT "how do we de-anonymize". That is super easy.
The problem is "how do we de-anonymize IN A WAY THAT LEAVES RESULTS OF RELEVANT STATISTICAL TESTS INVARIANT". Much harder. Remove a bunch of probably non-random points and make any logisitic regression on that dataset return the same result within some delta. There are trivial impossibility theorems. Really, really trivial. "Left as freshman exercise and if you can't some up with an example maybe step back and stop talking about this topic" trivial.
Now, add back consent. "how do we delete a subset of data that is likely not uniformly distributed with respect to risk variables and also de-anonymize the remaining subset of data.... AGAIN, IN A WAY THAT LEAVES RESULTS OF RELEVANT STATISTICAL TESTS INVARIANT". The position is untenable.
That is WAY harder that it sounds. When you add public disclosure of mismatch between results on public data and private (full) data things are even harder.
Specifically: suppose an independent researcher determines X is true via statistical test Y on the public dataset. Officials must respond, and say "True" or "False".
A small number of queries via this process can be used to infer data that makes re-identification possible. And there will be many more than a small number of queries, including from adversarial agents specifically designing queries to maximize information gain based on previous (public) revelations.
This is a game, in the game theory sense, in which the goal is to exploit the mathematical properties of statistical tests and social dynamics to force public revelations of binary assessments that yield information about individual cells or sets of cells and make re-id possible.
Have you actually studied this stuff, designed de-anonymization protocols, developed re-id threat models, thought about privacy budgets in public disclosure mandates, and at last put your reputation and your customer's privacy on the line when asserting that your statistical methods are resilient against attacks? Or are you just a jawboning know-it-all anonymous coward without any actual experience balancing privacy against public disclosure?
To be honest, all of your posts read like "just stop writing vulnerable software and then there will be no more hackers mmmkay?"
Facial recognition works perfectly fine on MRI face renders. You speak awfully confidently about a problem you seem to have very little experience in.
As with cryptography: Anyone can come up with suggestion that they themselves can't find an attack against.
FYI, high-res face MRI are just pictures of the face. So yes, if you share the picture of the face of someone, they may find your identity back, but it's not specific to MRI.
Dealing with any medical record that contains pictures of patient you should take additional steps to prevent identification, and face is the most sensitive part of the body as it can be most likely identified.
If tomorrow there are publicly available anonymised close-up pictures of cardiac MRI, it'd be extremely challenging to map them to their real owner with a 100% certainty but the discoveries it could lead to are very exciting (and the same with skin cancer pictures for example, or suspiciouses moles)
I only really hesitate to post full name and address but, frankly, I'm sure given my user handle it wouldn't be terribly hard for someone to find both of those.
So tell me, what harm have I done to myself by sharing this information?
None, and yet you still aren't willing to put your name behind the statement "I'm healthy as a fiddle and never use drugs except the occasional beer".
Also -- and this will apparently come as a surprise to you -- people who have STDs, abuse drugs/alcohol, or have genetic disorders in their family don't necessarily want that info shared on the internet. Shocking, right?!
Apparently pseudonymous internet griping is an essential right but participants in medical trials should -- post-facto -- have their entire medical history shared in a wy that's at significant risk of reidentification.
Oh genetic disorders, you didn't ask about that earlier. My family is a carrier for cystic fibrosis. Is that something that's supposed to be super terrible to share?
I'm not the GP, but frankly I think it's a little crazy how worried everyone is about being identified through round about methods in medical data. Google has a far better bead on exactly who I am than any doctor and, frankly, I'm far more worried about that advertising profile than I am about the "what if's" of my medical data being shared between hospitals and doctors (and potentially being leaked).
>My family...
>... I'm far more worried about that advertising profile than I am about the "what if's" of my medical data...
Again - yours. It's great that you're fine with your data being public, but as is evidenced by the way certain posts in this comment chain are being down/upvoted, clearly many other people share the opposite view and would prefer their data remain private.
AFAIK, the world hasn't ended and they aren't living in a dystopia because their medical information is so well documented. On the contrary it has served to stop the spread of genetic disorders and aid in medical research.
Certainly I personally don't have these hangups about privacy. I think it's weird that society has created them. Pre-HIPAA that wasn't the case and it wasn't exactly seen as major breaches of privacy for doctors/nurses to talk relatively openly about diseases their patients had. The sanctimony of medical data is a recent development practically unique to the US.
I'm not advocating that we broadcast to everyone details of every medical checkup. I am, however, saying that we shouldn't have nearly the level of restrictions we have around sharing medical information, particularly between medical institutions doing research.
Do you realize that people were murdered because of their medical history? Like mental illness and ADD were reasons for killing people in America just a couple of decades ago?
And currently, today there places where it is acceptable to drown your own babies if they demonstrate delayed development. Imagine your village elder coming to you and telling you that your baby needs to go in a bucket of water because he has club feet. That is happening. Today.
Citation needed.
You realize a couple of decades ago was 2000, right? The closest I can think of what you might be talking about is the eugenics programs ran up to the early 1960s. Even so, those weren't executions, but rather sterilizations. Barbaric, but not murder.
> And currently, today there places where it is acceptable to drown your own babies if they demonstrate delayed development.
And what part of making health recorders more accessible would influence that one way or another? Do you think it's the fact that heath records are hard to access that protects us from parents murdering their children? How are other nations in the EU managing to not kill all their kids for defects!
So, to be clear, your right to post random ass thoughts on the internet under a pseudonym is sacred... but it's totally reasonable to RENEGE on EXISTING privacy agreements with real world actual existing real human patients about intimate health records?
Just so long as we're clear about the enormous esteem you hold yourself in and the absolute contempt you have for others' privacy...
Second, demonstrating your willingness to describe yourself on a public forum is a lot different than someone expecting confidentiality when participating in medical trials. Most people do not have your level of comfort disclosing medical details. Furthermore, if a company were to break that trust, why would anybody participate in medical trials? The harm is not to you, it is to society.
This is irrelevant.
The conditional probability of my willingness to post this information, given your willingness to post this information, is equal to the the probability of my willingness to post this information. P(X) = P(X|Y). This means that just because you are okay with it, I'm not...and additional people being okay with it doesn't change my opinion regardless of harm.
Very happy for you, but if you did, you might have learned by now that people (and employers) can treat you very differently after they find out you're an addict or have a mental disorder, and it's best not to telegraph that information to everyone except close friends you can trust.
https://www.eeoc.gov/laws/guidance/your-employment-rights-in...
They are DEFINITELY not allowed to ask.
> "We are required to measure our progress toward having at least 7% of our workforce be individuals with disabilities. To do this, we must ask applicants and employees if they have a disability or have ever had a disability. Because a person may become disabled at any time, we ask all of our employees to update their information at least every five years.
Identifying yourself as an individual with a disability is voluntary, and we hope that you will choose to do so."
Just stop. This is embarrassing.
You're comfortable sharing it, that's cool! Is your neighbor? Your friend? Your cousin? Should they be obligated to divulge their own medical information just because you're fine divulging yours?
Address and name are not necessary.
- Person 30-35, living in California, has history of obesity, suffers from type II diabetes, under treatment of med A, and med B.
- first dose: Jan 2021
- second dose: March 2021
Medical outcome: ZZZ
What's not to get? OP said:
>... I strongly believe there is enough information to re-identify patients.
Their concern is that just dumping raw data like this into the public would be a massive violation of privacy for countless individuals.
It's for the benefit of patients, and this raw data is already available, but under a loose NDA for commercial partners and researchers...
And of course raw data like this could benefit patients, but keeping access to raw data limited also benefits patients. Sure, the data is readily available under a loose NDA, but it's available to people who have explicit knowledge of the requirements around handling sensitive, identifiable data. Average Joe's and Jane's do not have this knowledge, and bad actors just plain don't give a shit.
The fact that you poo-poo this question is telling. Preventing re-identification attacks is incredibly subtle and sometimes impossible. Removing "city and date of birth" is nowhere near sufficient.
> this raw data is already available, but under a loose NDA for commercial partners and researchers...
Yes. It's available to people who have their real identities tied to their access and can be sued into oblivion or possibly even prosecuted for wanton misuse of the data. Surely you see how this is different from throwing it on a public s3 bucket, right?
It reminds me the FDA saying they can't release the COVID-19 documents because reviewing 44'000 documents would take 50 years.
It didn't take 50 years to produce them...
Yes there is some effort needed to anonymize reasonably the data but it's not an impossible task. It's a question of motivation.
Here, clearly, the labs and administration don't really want to put efforts into that.
And, yes, I think public data on this sort of stuff is important and should be properly resourced.
I think:
Step 1. ASK PERMISSION before changing the way data is handled.
Step 2. Share full datasets more generously, but still gated by use and handling agreements. This probably means a private citizens without ethics research training and a supporting it dept can’t get a copy and peruse it on their personal laptop, but also that a truly enormous number of researchers would be able to access the data.
The gold standard of a CSV in a public s3 bucket shouldn’t be the enemy of good enough.
And no, this isn’t under a “loose” NDA. This would be covered by HIPAA, which tends to be un-subtle about violations.
0 - https://www.cs.princeton.edu/~arvindn/publications/de-anonym...
Isn't that provably untrue? https://bits.blogs.nytimes.com/2015/01/29/with-a-few-bits-of...
I'm all for putting pressure on everyone in Science to publish more raw data. This kind of data is likely more complicated because it's really hard if not impossible to anonymize the actual patient-level data. It still should be as accessible as possible to other scientists.
For example, when it comes to vaccine side-effects, I don't think there exists a true account for how common the side-effects really are. The most common way to report side-effects (VAERS, and similar national databases) are dismissed due to the self-reporting nature, local GPs frequently dismiss side-effects and tell people to just go home and take a Panadol with zero reporting going on (I had this happen to me - started experiencing severe chest pain 2 days post-Pfizer. Subsequently saw a cardiologist after months of pain and his comment to me was "I'm seeing young people like you daily and your cases are going widely underreported"), etc.
Likewise, when it comes to vaccine effectiveness, there are a million and one confounding variables from % of the population that already had natural immunity, covid variants, health, age, seasonality, societal lockdowns, isolation, etc.
Also, it's important I think for us to raise the bar to the highest possible standard when you're talking about a medical intervention that was forced under significant duress (loss of job, social stigma, public/medical shaming) on a substantial percentage of the world's population. We should not be content as a society to come within inches of worldwide medical authoritarianism without asking some seriously hard fucking questions and imposing the absolute strictest and highest possible scientific standards to justify why.
Orthogonal to the original conversation but have your cardiologist consultations yielded anything?
I also have chest pain for now more than 2 months after the second dose of the Biontech mRNA vaccine, but the tests revealed nothing abnormal. A few people in my entourage have been having similar symptoms but theirs has since receeded.
It doesn't help that search engine results for anything close to "Covid-19 Vaccine Chest Pain" are overran by both antivax conspiracy theorists and obvious propaganda. I couldn't find concrete information save from a few disparate accounts of similar conditions[1], despite the apparent frequency of those symptoms.
[1]: https://spectator.com.au/2021/11/my-post-vaccine-chest-pain-...
I received a diagnosis of pericarditis and had persistent tachycardia, mildest strain and my heart rate would shoot to 170 BPM not going below 85 while lying down. Now I'm back to normal and my resting heart rate is now 50-55 BPM.
We can simply look at infection, hospitalization and death in vaccinated and unvaccinated populations. If we properly match the populations, we can determine if the vaccine saves lives, and it turns out that they do save a lot of lives.
However, what is less clear today is whether there has been a net positive or negative effect of the vaccine for young healthy people. You can only come to that conclusion if you actually had high quality data and studies on vaccine side-effects, effectiveness in population groups stratified by age, health, etc.
I suspect that the vaccines, mandates, lockdowns, etc. have been a net negative for the overall health of young (<50), and healthy people, and the body of scientific evidence will support this position in the future. It's just cloudy today because it's wrapped up in politics...but the science will eventually win out.
Judging the risk/benefit ratio is the primary purpose of the regulatory agencies that approve vaccines. I don't see any reason to believe the claim that the vaccines are harmful for everyone below 50, that sounds quite outrageous to me. There have been adjustments based on new data for the vaccines a few times, e.g. younger people are generally recommended to be vaccinated with Biontech and not Moderna or AZ based on the side effects of these vaccines. That doesn't mean the risk/benefit ratio is bad there, it only means that we have vaccines with a more favorable profile for those age groups.
Another clear case of someone implying totally crazy things (younger folks without the vaccine would have been better without the vaccine) with absolutely NOTHING to support it.
However, collecting and analyzing that data would likely have eaten a percent or two into the eighty billion dollars Pfizer made last year, so I guess there's nothing we can do but trust the same authorities that brought us 'natural origin for sure', 'masks don't work', 'NNVTs don't mean anything any more', and 'Covid isn't airborne'.
If you bring lockdowns into the picture, you have to compare to what would have happened WITHOUT a lockdown as well, how many more deaths in hospitals, etc. The countries that tried this strategy have a very high excess death to compared to those that tried to limit human contacts (especially PRE vaccine).
Sorry but I can't distrust VAERS and then trust the COVID injection complication data added by the same people but now with financial incentives.
You don’t have to trust the drug co’s for that, we also have vaccine safety datalink system; so far the only notable side effect of the mRNA vaccines has been the myocarditis in younger people.
https://pubmed.ncbi.nlm.nih.gov/34477809/
Also your point about PCR testing is not accurate.
Source on that? Because all graphs I've seen show impossible to miss fluorescence around 35 cycles and up.
AFAICT VSD only does specific research at their own behest and currently don't have a section on COVID-19 vaccines.
Btw your link is dead.
I had pericarditis and some immediate reaction, my cardiologist thinks it was partially intravenously applied. Looking at the data on severe reactions from where I live I've been able to obviously tell that CDC must have used incredible criteria for their numbers. At least initially, I stopped caring when it eventually became clear to me that we do not really want to know how many are harmed.
And from the perspective of everyone involved I understand it, I too want this to be over, I too want this to be a safe magic bullet. But seems to me somewhere between 1:1000-10000 have significant heart issues from the Pfizer vaccine, but when we were rolling it out the numbers were claimed to be 1 in 230M.
It’s an editorial summarizing the first VSD report on Covid vaccine side effect research that I was describing (https://DOI.org/10.1001/jama.2021.15072)
I’m not going to argue that the US seemed to take longer and have worse communication about the myo/pericarditis issue than some other countries, but these things are being followed up on. The absolute timing I think is hard to discuss with a definite time frame
Here are some resources on PCR amplification:
- https://www.mcgill.ca/oss/article/covid-19-critical-thinking...
- https://www.thermofisher.com/us/en/home/life-science/cloning...
For serious risk quantification and causal analysis, we have things like the vaccine safety datalink, which links all electronic health records across a bunch of hospital systems, covering IIRC 3% of the US population. The UK has something similar I think. I think transparency in that system (VSD) could be better, but it has the same problem this thread is discussing, that anonymizing the records may be at odds with making the analysis reproducible.
The (small) controlled trial data can be combined with the (large) uncontrolled mass-population data, e.g. one can look at the larger dataset for corroboration of weak signals in the smaller dataset.
New controlled trials can be started, where further evidence is needed.
We may have loads of data from real people being treated, but good luck trying to analyse it in a scientific way.
We need effective vaccines, and we need to be able to trust them. Just treating people with a vaccine doesn't always give that trust and confidence.
(I hope I got my medical terminology right, I'm not a clinician, but have been close to covid-19 related medical research)
You can absolutely analyze real world mass vaccination data in a scientific way. Just treat pre-vaccination and post-vaccination as two different patients. The numbers are so large that other factors (environmental, genetics, etc) cancel out. So if your post-vaccination population has a rate of heart attacks 80% lower than the established norm, you know that is worth further investigation.
(Full disclosure: my family used to own a company that ran clinical trials for pharmaceutical companies)
Yes, you can do this. No, it's not the same.
With this kind of uncontrolled, longitudinal study, yes, some things are quasi-controlled: genetics, possibly lifestyle, other drugs, etc. Anything that plausibly doesn't change (much) in a single person from time A to time B.
Some things you can't control for, but still matter: changes in behavior due to the event itself. Placebo effect. Sample bias (e.g. the patients receiving treatment X were selected to receive treatment X because it was felt that they would benefit from it. This is subtle, but can really mislead. In the example of vaccination, imagine that the people most likely to vaccinate their young kids are also the ones most likely to keep their kids isolated at home, in a protective bubble...)
> So if your post-vaccination population has a rate of heart attacks 80% lower than the established norm, you know that is worth further investigation.
The key part of that is the last three words: worth further investigation. To get the final answer, you still need a controlled experiment.
So the moment we knew the vaccines worked, the science got harder. Well, that point also save an enormous amount of lives, so I don't think this is something to complain about. It's still possible to do a lot of good science now, even randomized controlled trials. They just can't compare to unvaccinated groups, but you can still do stuff like compare double-vaccinated with triple vaccinated.
You could easily find millions of potential study participants who don’t want the vaccine. The study wouldn’t be double-blind, but it would be single-blind and better than what they chose to do.
There is one study that listed negative effectiveness against Omicron for the vaccines. That study did not control for any confounding factors, it simply was not designed to measure this particular thing, it was focused on a different question.
There are some strong fundamental reasons why we would not expect the vaccine to be harmful with future variants. And even then we would still be able to detect if it was.
That's the point! That's why we need controlled trials. With an unvaccinated group. You know, the thing that you say we can't have, because it's unethical.
> ...without adjustments.
As usual, there's an XKCD for that:
> You cannot keep a working treatment from people just to do more science, that is deeply and fundamentally unethical.
I see this position put out there a lot and generally go unchallenged. For the record some people think the opposite: that it was unethical to unblind the placebo group so early. At the time you could estimate that continuing the placebo group might lead to ~60 unnecessary deaths from Covid-19 and a few hundred more serious cases. That is in the lucky scenario where the early success held. But in all scenarios, in return for a few dozen volunteers risking their lives, billions of people would benefit from randomized, placebo-controlled trial data which the NIH itself calls the "gold standard" as to "whether or not a treatment is safe and effective".
IMO the harms we have gone through from flying in the dark (people underhyping/overhyping the vaccines), multiplied by the number of people involved, make this a case where unblinding was unethical. It is very easy to imagine that far more life was lost in the general population from being in the dark than was saved by unblinding the volunteer study population.
(There's also the issue of whether or not it would have been practical to keep the participants blinded. I think that's a challenging topic of its own orthogonal to the ethical question.)
Are you talking about how they unblinded the study after efficacy results started coming in?
I disagree with both parts of this. Why not associate each patient with a number then log all data against that? Patient names and other identifying information should simply never be attached to trial-related data (except for in a well-protected lookup file with highly restricted access).
Making scientific data accessible only to other scientists is highly anti-scientific. Feynman, one of the greatest scientists of all time has many quotes around exactly this mindset:
"Have no respect whatsoever for authority; forget who said it and instead look at what he starts with, where he ends up, and ask yourself, 'Is it reasonable?' ... we will doom humanity for a long time to the chains of authority, confined to the limits of our present imagination. It has been done so many times before."
"Science is the belief in the ignorance of experts"
"Our freedom to doubt was born out of a struggle against authority in the early days of science. It was a very deep and strong struggle: permit us to question - to doubt - to not be sure. I think that it is important that we do not forget this struggle and thus perhaps lose what we have gained."
Because "patient name" isn't the only way a patient can be identified.
"Male, 27, admitted to Sunnybrook hospital for stitches to his forehead due to knife wound on January 20th 2022" is almost certainly uniquely identifying (actually it's made up so it probably identifies no one), is in the person's medical data for the trial, and clearly needs to either be redacted or in some other way separated from the other line that says "contracted HIV on January 19th 2022". Yet both are relevant when investigating causes of side effects.
With the ACA, the insurance coverage problem is gone. Pre-existing conditions can't be used as a reason to deny or change insurance rates.
The Fundamentalist problem exists, but seems like much less of an issue with LGBT acceptance being so much better now than it was in the 90s.
I just don't think that health information is so valuable that guarding it like a state secret is warranted. I'm ok with the notion of putting in basic safeguards like not attaching a patient name with the information but I don't see it has horrible if some system can infer absolute identity from that stuff.
After all, seems far more scary that my web browsing generates a far clearer picture of who I am than what you could glean from my medical records.
Having health information public and easy to gather would (potentially) be a significant boon to the study of health information. It would also make it a lot easier for us to make health information sharing systems so you don't have to fill out the same 500 forms every time you go to a different hospital or doctor.
When I say re-evaluate HIPAA, I don't mean "Hey, let's put everyone's name right next to every checkup and list it in a wikipedia like DB and email relatives about the results of every checkup". I mean "Hey, let's consider limiting the scope of the law and the penalties associated with it". It doesn't have to be a black and white thing.
That all seems like a worthy risk if it enables researchers to find that "Hey, looks like people prescribed medication X with medication Y tend to develop cancer Z way more frequently than the general public" or "Hey, looks like people with gene X respond way better to cancer treatment Y than the general public".
There will also be some risk that someone has a very unique combination of visits or treatments that does make them identifiable but that should be an incredibly low percentage and we need to weigh the pros/cons. I can also get a good idea of who has cancer by sitting on a bench outside a cancer center and watching who goes in.
The more data you release, the more difficult it is to ensure that subjects cannot be reidentified by combining the "anonymized" dataset with other publicly available data.
Important note: It's not just a case of putting in the work and being willing to share. There are impossibility results in this space. That is, there are reasonable formal threat models for which it is mathematically impossible to release any version of the dataset that is (a) useful (ie contains enough data to replicate findings) and (b) not subject to unacceptable levels of reidentification.
I am generally in favor of releasing data and code/spreadsheets, but anything involving patients becomes difficult quickly. There are also reasonable middle grounds between "no access for anyone" and "throw it in a public s3 bucket". E.g., making data more available to researchers and medical practitioners -- or even sufficiently interested & motivated members of the public -- but under strict data handling rules, enforceable audit trails, and legal consequences for being reckless with sensitive patient data.
Hell, odds are better than even that your employer made you sign a document that would allow them to sue you into oblivion for forwarding some sure-to-be-doomed product launch draft to the wrong person. Being at least that careful with fine-grained medical data on hundreds/thousands of patients isn't exactly unreasonable...
Drawing a hard line on 100% complete, anonymous, and consequence-free access to troves of patient data isn't a reasonable position.
[1] https://journals.plos.org/plosone/article/file?id=10.1371/jo...
And you can't simply strip out that data, because a lot of the subsequent analysis of effectiveness is dependent on it.
i am sure there will be a red herring in all the data that will cash a small enough shadow.
but this jibbering is about dragging science down to the level of far right ignorance, which is the onlymplace concervatives can win.
There is so little correspondence between reality and claims already- I think you’re making the mistake that the discourse is in good faith. It doesn’t seem to be to me.
I commend them on this approach, and can only hope that the vaccines are as safe and effective as we have been told.
[1] https://www.bmj.com/content/375/bmj.n2635 - "Covid-19: Researcher blows the whistle on data integrity issues in Pfizer’s vaccine trial"
What does this mean? That article is very specific, and is not sensational. He's been complaining about the BMJ for some time, and not just about these vaccines.
edit: open data is fine and good. The BMJ demanding it here is intentional sensationalism and part of a continuing pattern.
Incidentally, emember that antivax wasn't even a thing until The Lancet unleashed it on the world. Journals are often badly behaved in order to make headlines.
If you are them please reveal yourself and explain how you found nothing wrong with the original article. Most of these comments in the threads just seem like tired knee jerk anti big media / anti vax memes.
I would like safe and effective vaccines, and I would like to know that the vaccines I have been administered have really got the backing of truly scientific trials.
Surely releasing the trial data is a net positive for all the national populations that have paid for it.
And it doesn't help that there have been multiple leaks now regarding companies like Pfizer that have made them appear even more shady. Not to mention Pfizers multiple billion dollar lawsuits/fines in the past. I don't blame some people for being skeptical, nor do I blame them for feeling like they are forced to get something they aren't sure they can trust.
It's hard to just trust an organization that has already done multiple things in the past to make you not want to trust them.
The cases that feel even worse for me are the people who have had previous doses and are coming in for 2nd doses or boosters but have had previous unexplained side effects shortly after the doses. Like side effects that are beyond the "normal" ones. Many of these people have visited hospitals or their family doctors, had multiple tests, and get told that nothing is wrong, or that they are not sure what is wrong. And these kind of cases are often not counted an "adverse effect". I alone have had many people tell me these things. They often again feel forced to get the next dose because of pressure from their employer or government. And it just feels bad giving them the vaccine, because they don't sound anti-vaxx, they sound like they legitimately had a severe reaction to the vaccine. But often doctors and public health agencies will brush these incidents off and pretend like it's not because of the vaccine. They are dismissed and considered an anti-vaxxer making things up. And this really sucks. This is not the way medicine should work.
I heard that there was a public database in US for this sort of thing. But it has it’s own problems about data quality.
We have people who come in to get additional doses and they'll mention having some pretty serious symptoms after their last vaccine and we will be so surprised with how often we are told they went to both the hospital and their family doctor and got zero resolution and the doctor did not report it to public health.
In many cases we just seem them medically inelligble in these cases and tell them they need to go to their doctor and followup more to ensure they will be safe to get another dose.
And this is where I think a lot of the problems with vaccine mandates and pressure from employers can lie. You end up with people who feel like they have no choice to get the vaccine, they have a bad reaction to it, and then doctors treat them like they are lying or that it's no big deal. Then they go for their next dose feeling even more untrusting of science and get another vaccine.
We are ruining a lot of people's trust in science and medicine by how hard these things are pushed. Because of how much push back there is against anything negative with vaccines, it means doctors also often end up with this same bias. They are afraid to attribute any problem a patient has to the vaccine.
I have no problem with vaccines and have 3 doses currently. But I can empathize with these people who feel hesitant and I myself definitely have lost some faith over the last year of doing this work.
It really seems like a lot of scientific and medical standards have been tossed out the window because of politics.
We also have the vaccine safety datalink, which is intended for causal and quantitative studies of the things surfaced in VAERS, etc. it’s based on electronic health record integration across a network of hospitals covering I think about 3% of the US population. It is public as in publicly funded I guess, but not as in public access. You can read the first report here: https://doi.org/10.1001/jama.2021.15072
I am reasonably sure the UK has a similar setup, with some passive monitoring system and a much better health record integration story through NHS.
That is especially hard since there is simply no way to know for sure whether the vaccine will be "totally safe and effective for them": https://news.ycombinator.com/item?id=29206892 .
Right, and do you believe that a full release of the phase 3 trial data would convince them? I sure don't. Anti-vaxx is a fundamentally anti-expert stance, and most people have not logically convinced themselves into that position. Give them the full data set and I bet 9 out of 10 would ignore it, misinterpret it, read lies about it, or just assume that the data set is fake.
Transparency in science should include publishing raw data, as apart from re-creating the statistics in the report there is no other way the trial is repeatable.
Furthermore, it’s not obvious full transparency will help. You’re almost certainly not a medical researcher, so it's not clear if you could do anything useful with the full data set. And our experience with VAERS shows that bad faith actors absolutely can and will lie about it. I see absolutely no upside to releasing the full data set, and a lot of downside for those who volunteered for the trials.
It's unfortunate that the anti-vaxx idiots are misinterpreting or decontextualizing their content, but that's the anti-vaxx's fault and it's what they've always done.
Based on their original censored article, it seems like they're heavily invested to make things seem problematic, and carefully avoid any indication that the status quo is acceptable. I don't think it's purely misinterpretation when anti vax people become generally skeptical of govt and pharma corps.
Like could they not just say the vaccine is safe in these articles based on other evidence? They care about people being able to share their posts to Facebook, I can't understand why they aren't trying even a little to reduce the possibility for public misinterpretation.
It might seem obvious to you that the article would be misinterpreted or parts of it taken out of context by a lay public, but scientific publications are meant to be read mainly by other scientists, and this paper was written with that specific public in mind.
Even if this was intended as official public health communication, which it wasn't, public health communication is an absolute nightmare of a minefield. There's absolutely nothing you can ever say that won't be misinterpreted or misunderstood by some people, and the game is all about minimizing the potential for that. It's easier said than done. (And again, this paper was meant to be shared among expert colleagues in the field.)
Given that, a few disclaimers and limitations to help quell public misinterpretation is required. Not to do so is irresponsible.
Two things can be both be true:
- the vaccine is safe and effective
- parts of approach in running studies weren't as good as they should have been
Seems to perhaps apply here too.
The mistake in the AZ trial is something I consider extremely embarassing and stupid (assuming the published information about it is correct). It's the kind of mistake nobody that actually works in a lab should make, and I really question the competency of the people involved here. But in the end this meant that a part of the study got a half-dose of the vaccine. That makes the interpretation a bit more difficult, but it does not invalidate the trial.
So this editorial is kind of nutty. To give you an idea, this is kind of like asking banks to release customer account data to the public.
Are citizen biostatisticians going to verify the data and post about how everything checks out? That is quite naive seeing how the data from VAERS is being presented totally out of context in intellectually dishonest ways. Aside: the big problem with VAERS is that you do not have a denominator. Not to mention reporting biases, and the abysmal quality and assessment value of reports.
Even if you did have some cadre if motivated, ethical citizen biostatisticians any type of post hoc analysis would be suspect and epistemologically weaker than the original statistical analysis because it would have been designed with prior knowledge of the data.
Compare to the original study where the statistical analysis plan was prospectively designed and publicly disclosed BEFORE any patient enrollment or data collection.
Serious question: if the raw data could potentially undermine public trust in the US government, can it be classified as a matter of national security?
No, probably not. The government doesn't really have the authority to classify information generated by private entities like Pfizer barring specific legislation.
The closest you'd likely come is https://en.wikipedia.org/wiki/Invention_Secrecy_Act, but that allows (temporary) suppression of patents, not clinical trial data.
They can classify patents at will.
Pfizer is a defense contractor. It is pretty standard boilerplate that when you receive DoD funds the government can restrict access to information, even unrelated to the products or research directly attributable to that funding.
Not to say that I think the government would actually do that, just that they could.
Because it does. https://beta.clinicaltrials.gov/study/NCT04848584
> Several hundred million people have already received their vaccine, unaware that it's part of a study
Because it's not. Use under the EUA and subsequent approval is separate from and in parallel with continuing studies.
That is misleading at best. It absolutely is part of an ongoing clinical trial regardless us if it is authorized for emergency use.
How many people who were given the vaccine were told that it is part of an ongoing clinical trial, authorized for emergency use, and the data from the trial won’t be available for years? I suspect a very small percent.
Individuals receiving drug under an EUA are not part of a clinical trial, regardless of what you think of the drug.
It is for the people enrolled in an actual clinical trial.
It's not for people outside of the clinical trial using it under either the EUA or the subsequent full authorization.
> How many people who were given the vaccine were told that it is part of an ongoing clinical trial,
100% of those for whom that was true. Which is a very small percentage of the population who has received it.
> authorized for emergency use
100% of those for whom that was true, who aren't the same people.
See section e-1-A-ii here: https://www.law.cornell.edu/uscode/text/21/360bbb-3
(ii)Appropriate conditions designed to ensure that individuals to whom the product is administered are informed—
(I)that the Secretary has authorized the emergency use of the product;
(II)of the significant known and potential benefits and risks of such use, and of the extent to which such benefits and risks are unknown; and
(III)of the option to accept or refuse administration of the product, of the consequences, if any, of refusing administration of the product, and of the alternatives to the product that are available and of their benefits and risks.
So I am saying, what percentage of people who got the shot actually received this information? Since the trials are ongoing, that should have been communicated to 100% of people who received the vaccine as part of informed consent.That's very different than what you said before, and all of them who got it under the EUA.
> Since the trials are ongoing
The trials being ongoing is immaterial, it's required when under an EUA, not when the intervention is one for which trials are ongoing, which are different things.
Well, it is what I _meant_ and what I believe the op meant that you originally responded to.
> The trials being ongoing is immaterial
No, it is not immaterial, because people have to be notified of “the significant known and potential benefits and risks of such use, and of the extent to which such benefits and risks are unknown”. The fact that the clinical trial is ongoing means that many of those benefits and risks are, in fact, unknown.
You can spin this however you like, but my point is that most people who are receiving and have received shots have NO IDEA that clinical trials are ongoing and the actual data has not been made available to the public (even to their doctors who are likely recommending it). This is the reality whether you would like to believe it or not.
This is one of your many points of confusion; it is completely normal for even long approved interventions to have many clinical trials in varying stages. There are currently 92 trials for the Pfizer vaccine listed on clinicaltrials.gov, in stages from “not yet recruiting” to “completed”.
And also 1,761 between “Not Yet Recruiting” and “Active” for the notorious untested drug aspirin.
> many of those benefits and risks are, in fact, unknown
That's kind of expected with an EUA more than a fully approved intervention (but even for the latter it's always the case), and why disclosure of potential unlnown risks is required with EUAs.
You unsupported assertion that this requirement is routinely ignored would be a problem if it was true, but you offer exactly zero reason to believe it is true, surrounded by a lot of irrelevancies that show you don't understand the context.
Long-term studies are common. We still study Tylenol to this day, for example, to monitor for things like "does it interact with newly approved drugs".
https://www.hhs.gov/hipaa/for-professionals/privacy/special-...
I also worked on medical devices and respective medical advisory boards. Getting people to share their data is just a question away. If you fail at that, I doubt any further inquiry would net any benefit.
I presume this would actually be the solution if the "raw data" people were making good faith requests.
Is there any reason why this wouldn't be viable?
When it comes to spread the vaccines may be even making the situation worse as the infected vaccinated people spread virus at the same level as infected unvaccinated while having less symptoms for the vaccinated means higher chances of them walking around spreading it.
>there were very troubling results from the clinical trials that were swept under the rug
this is why the clinical trials aren't going to conclude until few years down the road as operating under EUA is basically get-out-of-jail-free card for the BigPharma here.
We have zero idea what the long term side effects may be. I hope it is zero. But if it’s not there are going to be some very unfortunate people.
75 years, actually.
I don't mind trusting what someone says, but I need to be able to verify it.
They deserved to have that both mocked and slapped down. It was an insult to both our intelligence and our right to know the truth about what they want us to take.
> Indeed, Doshi has history of playing footsie with the antivaccine movement, amplifying antivaccine conspiracy theories, downplaying the severity of influenza and thus feeding antivaccine narratives, using sleight-of-hand to downplay the effectiveness of flu vaccines, and generally playing the role of a false skeptic with respect to vaccines. At this point, I can’t help but note that Doshi also once signed a petition “questioning” whether HIV causes AIDS.
How about you argue with the points made in the article? Do you disagree with anything in the article? Is their dreaded "misinformation" in it. If so point it out and ideally provide some evidence.
I am pro transparency especially when it comes to a global pandemic. I've always been that way. I'm not going to change my opinion because some other person has the same opinion. I don't give a rats ass who wrote the article. It could have been Santa Claus. I like to read articles and do my own research and form my own opinions. People like you who want to try to discredit people so you can silence debate are the biggest problem in this pandemic.
Oh, and you know who else got AIDS and HIV completely wrong? Dr. Fauci.
Absolutely, yes. I believe there to be an acute public health interest in not letting antivax nuts run wild on raw data like they do with VAERS, despite clear disclaimers not to do so. Additionally, as a participant in a couple of trials, I see a significant benefit to careful review before the release of data that might inadvertently reveal my identity, conditions, medical history, etc.
Conspiracy theorists usually do go nuts with this stuff, but usually there's no need for an organized "conspiracy" when class or professional interests are at stake.
For example, that's one very clearly stated reason why US government officials running scientific agencies shut down scientific discussion of the lab leak theory, and why most scientists went along with it: because of the risk that the public backlash might harm their funding - the same sort of backlash that led to restrictions on gain-of-function research in the first place, that they ended up working around.
> ‘I share your view that a swift convening of experts in a confidence inspiring framework (WHO seems really the only option) is needed, or the voices of conspiracy will quickly dominate, doing great potential harm to science and international harmony…’”
People quite naturally and regularly act in a way that appears coordinated without any actual coordination.