How Anne Wojcicki’s 23andMe Will Mine Its Giant DNA Database
forbes.com
forbes.com
We often see complaints about big companies selling their users' personal data, but in this case the decision lies with each individual who shares his DNA sequence. Do you believe it is ethical to share your relatives' personal data without their consent?
They don't ask me, I don't approve that, but still Facebook gets my data. Same applies for Viber, Signal, etc. of course.
https://support.signal.org/hc/en-us/articles/360007061452-Do...
(Of course sharing /w the Facebook app is probably a completely different story.)
They used the faces in an art display.
I’ll see if I can find it but I’m not confident—
Edit: found it easier than I thought:
https://www.cnn.com/2013/09/04/tech/innovation/dna-face-scul...
This is a privacy matter, I think. It seems haphazard to allow this kind of thing wantonly...
For instance, your DNA used to be private in the past simply because no one could do anything with it even if they would get it. So you benefited from this privacy aspect by default. In the next 20 years your DNA will probably be in multiple databases somewhere even if you never offered it to companies.
Similarly, in the past you may have benefited from the privacy of your own home. You could say or do anything you wanted and that would be kept private (for the most part). Now, with all the "always-on" smartphones and smart home devices and surveillance cameras, everything you do or say in your home will be on someone else's server, which can be data mined, sold to third-parties, requested by various law enforcement agencies, and stolen by cybercriminals.
It seems to me that from your point of view and with enough technology advances we'll have 0% privacy in the future. Everything around us will listen to us and watch us, and then others, with who you may have never interacted, will also get to see and analyze all of that data.
So the technology could enable all of this -- but the question is should we let it? Privacy is a human right for a very good reason -- abuses against someone's private life can lead to all sorts of nefarious things against that person, whether it's something as "benign" as increasing your insurance rates to not offering your free/cheap healthcare because you "live too dangerously" or malicious actors and government agents using it to destroy your life for profit or personal vendettas.
We should fight it as possible, but we'll see how far that gets us.
You say "should we let it" like that wouldn't mean getting way more draconian about policing information. How do you prevent people from sharing their own DNA? How do you keep people from reading and remembering it?
It's not about the ability to look at something, face image or DNA. It's about ability to remember that information and use it.
Some argue that they want ownership of information that concern themselves, and right to deny others from handling that without explicit consent. Others argue that collection is alright but sharing (with or without exceptions) isn't. Others disagree completely and say that information belongs to whoever holds it, as it's immoral to deprive person of rights to have memories. It's also arguable that privacy laws may apply differently between persons and companies. Or between natural and artificial memory banks. Or that different rules should apply based on volume of data stored, like number of persons involved. I've read many different opinions on this matter. Lean towards some, but haven't firmly decided on any.
I don't think this will be fully resolved until someone would invent eidetic memory drugs or implants or something like that, and we'll have to erase the current boundary between humans and machines.
Famous example (albeit of a smaller significance, since the data of users shared was public for all their friends) - the whole Facebook/Cambridge Analytica scandal. Users were giving up their data using their survey app, but they ended up sharing the profile info of their friends as well without their permission.
Here's a service. Pay to get your panel/exome/whole genome sequenced, put your data on gcp/aws/azure, user open source to do variant analysis and pay a specialist for their computer aided analysis.
Guess what, you own your data end to end. I imagine the above could be had for less than $2k. Some startup will write the software stack to do this and simply sell licenses for software you run anywhere you want.
Sell it for use in embryo selection, employee profiling, to governments.
The scientific community has already created all the tough parts of the software stack, and everything has been commoditized already, there's just not much of a customer base.
The customers that it's really really useful for are those with rare phenotypes that want to learn more, and they can't find a doctor with expertise, and there's no lab specialized in genetic testing for that.
However, those are also the cases where the off the shelf software/sequencing doesn't always give the right answer because it's not a simple small change, and then you need specialists both on the biology and bioinformatics side to really make it useful.
> Reports of genetic testing must be retained for at least 10 years. Electronic records are acceptable. Specific regulations for specimen retention are not proposed, but each laboratory must have a written policy defining its own specimen retention policy.
Data retention is the next big wave and these companies would do well to clue in. There is zero chance they will find anything medically useful with their database and a ton of outstabding liability once people realize how it can be abused.
"CLIA certification and CAP accreditation 23andMe laboratory testing is done in U.S. laboratories certified to meet CLIA (Clinical Laboratory Improvement Amendments of 1988) standards, including qualifications for individuals performing testing and other standards to ensure the accuracy and reliability of results. The laboratory is also accredited by the College of American Pathologists (CAP), which has served as a model for various federal, state, and private laboratory accreditation programs throughout the world."
FDA authorizations for 23andMe's personal genome service are available online, e.g. [2] for Alzheimer's disease risk reporting based on the E4 variant of the APOE gene.
The company also offers ancestry reports, which are not clinical and thus covered by CLIA. But medical claims in 23andMe's health reports do comply with CLIA and other regulations.
1. https://medical.23andme.com/dna-kits/#clia
2. https://www.accessdata.fda.gov/cdrh_docs/pdf16/DEN160026.pdf
https://www.openhumans.org/member/iandanforth/
I want this data out there to help any and all researchers. The more free, public data exists the easier it is for researchers without GSK levels of cash to make discoveries and contribute back to the community. Just like OSS, someone has to be willing to give away something that has historically been sold. I'm willing to do that and I hope you will as well.
Valid concerns, but I think medicine does need vast amounts of genetic data available to actually make some progress. Medical progress is not advancing fast enough by some very real metrics (5 year cancer survival rate over 40 years is not at ALL impressive when you factor in diagnoses being made earlier).
Ownership of this data is important, but if genetic data was more freely shared I think medical progress would benefit.
How does more genetic data equate to better progress with medicine? Are we testing medicine against specific genome sequences, now? If so, how are we doing that without the source host[s] to test against?
>Medical progress is not advancing fast enough by some very real metrics...
Again, I'm not seeing how these two equate, whatsoever.
For what you're talking about the genomes would have to be reproduced, such as the markers that are the precursors for breastcancer. Then, the tissues, themselves, would have to be reproduced and then you'd have an effective field with which to test new medicines against (unless you just use the people with the markers to test).
How you're getting from a mapped sequence to better medicine is... ...we're simply not there, yet, technologically, I believe. Unless you know something I don't? (Which could very well be the case, admittedly, but I doubt we're at the stage of computer models for genome engineering, tissue growth, cellular division, etc. all in one.)
Take heart disease. It has a significant but complex genetic component. Many genetic variants each contribute a small amount to risk for heart disease. If a given person has many small risk variants, the sum total risk -- often called "polygenic score" -- can be relatively high.
People in the top 8% of polygenic scores had a 3x higher risk for heart disease than the general population [1][2]. Through techniques like polygenic scoring, large genetic datasets enable uniquely early detection of high risk for the world's leading cause of death.
[1] https://www.nature.com/articles/s41588-018-0183-z
[2] https://www.vox.com/science-and-health/2018/8/24/17759772/ge...
[0] - https://twitter.com/cecilejanssens/status/103135930540723404...
Also, you're posing prediction and targeted treatment but you haven't posited how Bob's mapped genome sitting in 23andme will be used for medical treatment.
As we know, genes are not an emphatic, "this will happen to you," but an increase in likelihood; which still doesn't translate to any emphatic treatments from the genes, themselves, yeah?
Janssens seems less skeptical of 23andMe's paper on polygenic score for type 2 diabetes [1][2], which -- interestingly -- positively cites the Khera 2018 paper on polygenic score for heart disease that she critiqued. Some researchers are skeptical, but the medical community generally seems to consider polygenic scores promising for tests [3][4].
> you haven't posited how Bob's mapped genome sitting in 23andme will be used for medical treatment.
Early intervention. Polygenic scores could be used for medical treatment by motivating earlier intervention. That could include stronger recommendations for better diet and exercise, closer monitoring programs, or more precise prescriptions. That, in turn, could reduce disease burden.
[1] https://twitter.com/cecilejanssens/status/113707970323438797...
[2] https://permalinks.23andme.com/pdf/23_19-Type2Diabetes_March...
[3] https://twitter.com/EricTopol/status/1129780543434964993
[4] https://journals.plos.org/plosmedicine/article?id=10.1371/jo...
and
>...could include stronger recommendations for better diet and exercise, closer monitoring programs, or more precise prescriptions...
and
>...could reduce disease burden...
This is where the problem delineates for me: We're being massively assumptive in moving from "could" to "is" and "will".
I will, generally, concede the could portion but to assert that it is emphatically happening or going to happen is still far from fruition and to label this science as such, just yet, is overreaching and giving false hope where none should really be given because, then, you'll taint it's benefits with the drawbacks.
Remember: Anonymised data (e.g.: 23andme) only allows a survey of what's relatively known or can be inferred from the anonymised dataset.
To arrive at what you're suggesting, it would have to move into a different realm (I believe), like UK BioBank or GEDMatch but, even then, we're still basing things on speculative science - gambles of percentages that aren't, emphatically, true or false but a kind of "maybe, kind of, sort of, in a way, definitely could or defintely could not" muddied waters.
That, to me, is a far stretch from saying that the data in 23andme is - actually - helping medicine; which I believe is what the OC I replied to emphatically said.
Why not fix both of those things first before feeding our personal data into a fundamentally broken system?
I know we are all good people, but most of novel science happened with methods that we would find totally unacceptable today, from lobotomizing people, to deafferenting cats, to sending dogs to space. If 23andme didnt have established a large user base, it would be impossible to get a comparably big database given today's regulations (and climate).
And it 's not like they lock-in your data, you can export them to contribute to other studies if possible.
All UK Biobank participants were tested for blood pressure, bone mineral density, grip strength, BMI, etc. The last 200,000 participants underwent detailed tests of cognitive function [1].
23andMe asks customers for health information, but self-reports are not usually as reliable as the clinician-administered tests done in UK Biobank.
What I mean is you pay 23andme money, they showcase your genetics which, as I understand it, include lineage, risk of disease, etc. 23andme isn't akin to a blood panel in my mind. At least not at this point since they sell on other marketing aspects that are far more consumer oriented.
> I know we are all good people, but most of novel science happened with methods that we would find totally unacceptable today, from lobotomizing people, to deafferenting cats, to sending dogs to space. If 23andme didnt have established a large user base, it would be impossible to get a comparably big database given today's regulations (and climate).
I don't agree with this. You're conflating lobotomy and, again, a for profit company that is collecting a massive database of DNA as we speak and all of the use cases are not yet defined. There is a significant difference between scientific research and running a for profit company. If we can't agree the motives between those two things are different then that is the clear disconnect. Swap out your analogy with Google or Facebook. Are we all glad they have established a large user base? I'm not. And just like those companies it may be that you may think 23andme having your data today is a good thing but that could change quickly when your data is sold for different intentions later on. Do you agree you may not always hold that perspective? What if 23andme starts to sell data to insurance companies after they lobby away having to cover preexisting conditions and then buy your user data to deny you coverage?
As for the privacy concerns i have a bit different views. I think it s debateable whether DNA is intellectual property of its owner and whether it could be protected as such. Very few humans can claim that their DNA information is the output of their labor, and it is very highly shared among humans, so there should be a discussion to what extent genetic information should be subject to fair use. Having widely available open datasets would reduce the value of these big centralized databases.
Our government and insurance systems are shitty, and until they are fixed, disease and aging won't be our number 1 priority.
Even individually, our DNA is probably worth a bundle to many companies. To me, my DNA is as near priceless as something can get. I'm glad I didn't pay someone to take it.
The profitable, yet dystopian uses of this information are potentially limitless (see the most recent thread about how the Nazis used IBM technology), but the regulations are negligible to non-existent. I'm more than happy to not be one of the victims of these uses for whom laws and regulations are made after-the-fact, if they're made at all.
Health insurance is not important here (UK), and life insurance is much less of an important concern to me as I will not have children. But I can see why if that is not the case, you would be more concerned.
I do have some concerns about it being used by the police, though public availability is unlikely to make a difference there. I've never committed any crimes, but the police are known for using scientifically and statistically invalid methods to prosecute.
Also, in 10-20 years, what if new threats arise?
If it's valid trait that employer (or insurance) finds actually having negative impact - is it really a discrimination (other than you can't really say whether it's expressed).
[0] https://en.wikipedia.org/wiki/Genetic_Information_Nondiscrim...
Same way you get to know your KPI's and other assessments?
I guess that act is good after all. No one can tell you've got an actual disease from these tests, even if on big scale the correlation most obviously is there.
Don’t I leave a trail of DNA everywhere I go anyway, from shedding dead skin cells?
There are very serious reasons to keep your genetic information private [0]. Unlike a photo, genetic data reveals sensitive medical information and more...
(+): Well if someone wants to engineer a deadly virus that targets specifically me, then what can i say , that's at least a memorable way to die. But they can do that from a piece of my hair too.
It is conceivable that this would require massive amounts of money to lobby politicians and inform the public of an issues that most people don't give a shit about (Hell, people are so scientifically averse that vaccination is still a significant issues in 2019).
What makes this even more challenging is that there are parties whose interest is to exploit that information for their personal gain rather than helping the individuals with that DNA (e.g. insurance companies, government agencies, etc.). Those parties likely have access to more money and are more capable of lobbying and propagandizing for their interests.
Humans are ballooning as a species and hurting the planet and causing other species to go extinct... and you want to eliminate disease and aging from said species?
https://www.populationpyramid.net/static/population-projecti...
Or to be specific, Africa and West Asia, due to a lack of access to contraception. That can be solved by redirecting 100% of aid towards education, contraception and abortion for women, and away from food aid.
I understand the point of view that preserving species is a moral good, but it is just one of many possibilities and could very well be wrong.
DNA SNPs are one thing, but genetic expression (phenotype) is immensely complex due to the number of variables. For example, mammalian immune systems are highly redundant across the body and continually signal with cytokines and chemokines to manipulate the immune response that is internally/externally environment dependent. Our bodies continually work to remove/detox OR sequester pathogens, toxins, metals, etc. This usually results in isolated and/or systemic inflammation across the body. Many of the pathogens manifest with overlapping symptoms making it hard to isolate without better diagnostic tools. For example, my two strains of Bartonella cause inflammation in many similar areas that Lyme bacteria do and Babesia has another similar set. It is hard to isolate those symptoms from the mold toxins and metal (aluminum, arsenic, lead) buildup that I also have. Functional doctors believe the HLA SNPs play a role in accumulation rather than removal. Enter saunas/sweat tents/etc that many cultures used for centuries (lymphatic detox).
Enter public funding to invent better and cheaper measuring equipment where each of us owns the data. Privacy protection is also important. We live on a planet with immense biodiversity, albeit shrinking daily due to human activities, where we coexist with bacteria, pathogens, molecules, industrial toxins, etc -- It's time we start to learn more about everything ... gathering data in a coherent schema and applying ML can certainly help.
That is why there are a number of studies that list various pathogens in brain deposits for Alzheimers, MS, Giant Cell, etc -- Feel free to google and research many of the terms above and reach your own conclusions :-)
(Of course the data doesn't have to be accurate or the analysis correct to suffer.)
Doesn't seem a big enough concern, how likely I'm going to win lottery?
How about the insurance company? Forced them to lose money?
They won't hire me because my ancestry? Well they are stupid then, it's their lost.
manual grunt work will be replaced by automation anyway regardless of you develop chronic illness or not.
You don't own your fingerprint either. Or your height or weight.
By the same argument all companies don't own their code or IP as soon as it's in someone else's hands.
https://www.bostonglobe.com/metro/2018/01/12/man-found-guilt...
Child pornography is information that spreads. I believe society is still healthier if that kind of information is never produced or spread in any form.
And besides, 23andme's research data is anonymized.
Although both 23andme and GEDmatch match up relatives, I don't think they share your actual genetic sequence with your matches. 23andme's research and data publications are anonymized.
2. There is no such thing as anonymous genetic data once you move past extremely small SNP panels. Even if the data is pooled.
The solution is to ban discrimination. Banning a technology that has the power to actually make the world a better place through better medical interventions is not a trade off I want to make in the name of ending genetic discrimination (I write this as an African American).
The solution is a single-payer system, preferably state-run health care.
Your DNA happens to be at the scene of a crime (you didn't commit). Police do a DNA dragnet. Now good luck getting off that case.
Now imagine a health-reporting institution, used by employers etc. Similar issues will happen - mistakes, failures to update. You get a 'high health risk' rating and can never get insurance again? That's an issue.
You leave a DNA trail everywhere you go. There is no way to keep it private.
Also false positives for DNA matches - overconfidence in results from DNA labs due to popular culture thinking DNA match = 100% guilty combined with numerous controversies over police forensics in recent years, Houston off the top of my head had to throw out every case that their police lab did testing for a range of years due to lab tech incompetence.
Same with ubiome.
Do I care that a company can make a profit? No.. I prefer bsd/mit licease as well for this reason.
This is the most responsible way to use the data.
The way this model works is that a pharmaceutical companies have a question that requires a genetic database to answer. They then pay 23andMe for the results of these tests.
23andMe does not have to give up any customer data to any third party and can at the same time build a solid business model around using their data for positive medical advancements.
How?
From the opening sentence, I already don't like her.
And now this. I wonder if that employee was just not very attentive to what was actually going on, or perhaps the company leadership deceived the rest of the company? Or maybe this is just yet another example of a once-noble company inevitably succumbing to the allure of "growth at all costs".
- Medical legitimacy? “Effectively what you have is a technology that is neither that helpful [nor] that harmful (...) Dr. Jonathan Berg, a clinical geneticist at the University of North Carolina at Chapel Hill."
- PR Signaling? “It was so unacceptable to her as a compassionate human,” says Ashley Dombkowski, (...) “She is undeterred by massive, worthwhile problems.”
- Personal Branding? Cue the Jobs-esque clad-in-black shots.
- 'Certainty' in the face of material uncertainty? "By October 2013, 23andMe was in talks with Target and Wojcicki was pushing hard to enter stores before the holidays (...) Emily Drabant Conley, recalls Wojcicki’s certainty when an exec thought it was impossible to meet the time line: “Anne was like, ‘This is a company that was founded on impossible.’” But in the end, impossible won. It would take three more years for 23andMe to get onto Target’s shelves. "
... so where have seen seen this before?
I'm glad that Forbes is taking a slightly more skeptical tone this time, at least calling out some of the issues and not taking PR at face value.
I actually hope they are successful, though I'll never participate due to obvious privacy concerns.