$200 retail whole genome sequencing now available
us.dantelabs.com
us.dantelabs.com
More worry is that they sent out used kits to a number of customers!
https://www.cnbc.com/2018/06/15/dante-labs-dna-testing-compa...
But would reviewers detect if the results were not really theirs and just from someone that matched on a few data points either from their forms or from a cheaper test?
You can find more reviews on our Facebook page
Maybe when we get a clear answer, I'll consider checking my genome. But not before.
> You may withdraw your consent to participate in Dante Labs Research at any time by contacting Dante Labs at the email address: contact@dantelabs.com. Dante Labs will not include your Genetic Information or Self-Reported Information in studies that start more than 30 days after you withdraw (it may take up to 30 days to withdraw your information after you withdraw your consent). Any research involving your data that has already been performed or published prior to your withdrawal from Dante Labs Research will not be reversed, undone, or withdrawn.
I'm not quite sure what to think.
Dante Labs will not sell, lease, or rent your individual-level information to any third party or to a third party for research purposes without your explicit consent.
So they will sell your data unless you specifically look for, find, and check the opt out. What happens to your data after that is out of Dante Labs control, so any other assurances they make are irrelevant.
My assumption, since this promo is for a few days only, is that it's like any sale; they're doing it to build their brand and attract new customers.
1 - in the event of a sale, your data may be among the assets transfered
2 - this policy is updated from time to time and comes into effect at the time of posting, and maybe we'll notify you ahead of time
Could they still sell it as marketing data?
The GDPR does, actually. The data can only be transferred to a GDPR-compliant company, and the user must agree to the transfer.
[0]: http://www.completegenomics.com/ [1]: http://www.rdrmanac.com/
That’s much higher than I’ve seen reported for prior Complete/BGI datasets. Recently lower error rates have been reported, but I’m surprised to see them this low.
If this is true, it’s comparable to the error rates on Illuminas older instruments (the GA1s and 2s at least, if not Hiseqs).
Q20 is one error in 100. Q30, one error in 1000. I think Illumina average error rate (Novaseq) is >Q30 now.
Veritas Genetics were also offering a 200USD genome a while back. I think that was Illumina sequencing, but it was a very limited time offer.
Yes, that was on Monday Nov 19th. Veritas offered 1,000 genomes (Illuminca sequencing) at USD 200. They sold out in 6 hours.
* Disclaimer: I work for Veritas.
[0] https://www.genome.gov/27541954/dna-sequencing-costs-data/
[1] https://techcrunch.com/2017/01/10/illumina-wants-to-sequence...
The only way this company stays alive without significant venture funding (which it appears they don't have) is through partnering with pharmas. From reading their privacy policy, it looks like they are de-identifying genotypic and phenotypic data from users. This, at scale, creates a nice population based genetics study so that if I'm a pharma, I can look at a specific biomarker and all data associated with it. That helps me build drugs against certain genes, diseases, etc.
So yes, you are the still the product and yes, your data is being resold.
At one point, we thought that simply replacing a user's IP with a unique ID was enough to anonymise.
At one point, a blur or pixelating filter on a photo was sufficient to anonymise it.
At one point, Monero transactions were considered anonymous.
That means that by definition, anonymizing genomes is impossible. If someone tells you they will 'anonymize' your genome data, run away, they don't know what they are doing.
http://41j.com/blog/2018/11/illumina-consumables-are-90-prof...
Their overall gross margin is >60%. I suspect that instrument margins are actually also quite high, but they factor in R&D spending on instrument costs.
However, Dante labs appear to be using BGI, not Illumina for sequencing.
Also, in their About Us section they have every technology listed under the sun (Illumina, PacBio, Thermo, BGI) so I think it's just whatever they can get their hands on based on which labs have open runs. Seems quite questionable from a data integrity perspective if you're trying to do population based genetics. The quality between sequencers varies dramatically.
The BGI (MGI) have been pushing out new instruments, reportedly with lower error rates (and much cheaper than Illumina).
They’ve also been building out service centers outside China, and I suspect this is where the sequencing being done by this service is located.
Universities do sometimes offer sequencing as a service, but I suspect not at as low a price point as the BGI, even for those universities with attached genome centers.
We go up against universities all the time in our competitive deals. They're generally always cheaper since everything is subsidized however the turnaround times are usually significant and the quality is not entirely repeatable. The cost here looks heavily subsidized just as a way to spark interest.
Even so, Dante is currently definitely operating at a loss. That’s not a problem though since their business model probably isn’t really consumer genetics. Likely, their actual business model (like that of every major player in the field) is gathering large panels of sequence data to either analyse themselves (and then sell the insight to pharmaceutical companies) or to sell (pseudonymised or aggregate) access to companies. So, in reality, they are preparing large-scale biomedical studies. And not only are they not paying study participants (their ostensible “customers”), they are actually getting the study participants to pay for their participation.
(Note that there’s nothing necessarily nefarious about this whole business. It just explains why they can offer their product to consumers “at a loss”. In reality it’s simply an investment.)
Wasn't Dubai trying to collect DNA on every citizen and visitor to their country?
Not saying this company is in anyway connected to Dubai but there are any number of reasons why a company might want to amass a large sampling of DNA and, although not a lawyer, I can think of some plausible ways to use that data too.
If they burn through any VC money they have amassing a database, then file bankruptcy or sell off their assets to satisfy debts before dissolving... "hey other company, that's totally not us or our investors, we gotta pay our taxes so we can dissolve, wanna buy some genomes wink wink?"
If you're really interested in DIYBIO there are subreddits and many tuts out there, but fair warning: getting started is easy, but the devil is in the details a heck of a lot more than learning to code.
[1] https://www.reddit.com/r/promethease/comments/9866da/dante_l...
>Receive actionable insights based on solid genetic and clinical evidence for your and your doctor.
Some preliminary ideas:
- misidentify yourself (eg get a few friends and submit each other's samples)
- mix your DNA with a mammalian organism (something like a mouse). It will get flagged as a bad sample by them probably, due to poor alignment metrics and too many variants, but with the knowledge of which DNA you put in you can do a dual alignment to eg. the mouse and human genome and purify the reads that came from you (there are published methods to do this kind of thing). Not sure how you would get an equimolar mix of mouse and your own DNA, but for $200 you could submit multiple ratios I guess.
- enzymatically modify your DNA. There are some enzymes which cause characteristic mutations. They will turn up as real variants in their pipeline, but you can use the sequence context to filter them out. Kind of like chemical code obfuscation. This would mean extracting your own DNA and treating it, probably not that feasible.
The first one is the best because if enough people did it, it would create enough doubt in the accuracy of the stored datasets that they wouldn’t be worth anything to anyone except the person who owns the DNA. The problem is, as this becomes more popular more people related to you are going to get it done, meaning they will be able to automatically fit you into an identity based on who else had it done, their known identities and just the similarities between their DNA and yours. Trying to maintain anonymity only works if the data set is small. It looks like there is a market for a completely discrete secure dna sequencing company that doesn’t keep a copy of your genome after the fact and sends you a copy that is encrypted. Also it might be feasible to have some kind of kit that divides up the chromosomes into groups to be sent to different labs.
It can’t be that difficult and expensive to do this. Why not just wait until there is a kit you can buy to do it directly from your PC if you are concerned with privacy?
The problem with 1 is that there may be significant downsides to misidentifying your genome eg you swapped with someone who commits a crime.
I wrote a summary of genetic information discussing this very idea. See the PDF under "What is Genetic Info?", read the Phylogenetics section:
geneinfosec.com
Also note that DNA is readily identifiable, so misidentifying yourself does not work.
[1] https://nanoporetech.com/applications/dna-nanopore-sequencin...
The reads lengths might be fascinating with Nick Loman's group reporting upto an Mb in size. But the error rates as well matter. Unfortunately nanopore errors are not easily mitigable, whereas PacBio's are.
The one place where nanopore has them beat is the sheer portability and ready run time.
Strains involved in 2015 ebola outbreak were being sequenced real time!
Looking at things like this - https://phys.org/news/2013-03-easy-identity-cell.html
https://33bits.wordpress.com/about/
etc
Having DNA and just looking at the traits that person has will reduce the population down to a pretty manageable number. Any other bits - location, gender, etc will be enough to make an educated guess to an almost certain one. This will not become more difficult to do in the future.
[1] http://www.pnas.org/content/early/2017/08/29/1711125114
[2] https://www.technologyreview.com/s/608813/does-your-genome-p...
I used 23andme years ago, using a fake name on the website and a kit ordered by someone else (a friend ordered kits). I should see if I can use their GDPR process to purge it before any relatives give it a shot and start a scandal about the presence of John Doe in the family tree.
[0] https://globalnews.ca/news/4171752/golden-state-killer-dna-g...
Right now, it's the wild west.
In what regard? I’ve not heard of any privacy related issue with 23andme, and I follow the field quite closely (professional interest). On the contrary, 23andme has been known to uphold customer privacy in the face of government requests [1]. The FTC is investigating 23andme and other companies (which is a good thing), but there is no indication that they’ve found any violation.
There’s been a lot of work with homeomorphic encryption, but I think that the new avenue of research using sketches to provide a sense of anonymity while preserving comparability.
To be fair, for most things the VCF is completely sufficient (and in fact most people won’t care even about that). It just feels cooler to be in control of the raw data (and personally if I end up using a sequencing service, I would want to perform my analysis; but this is obviously irrelevant for 99.99% of users).
What we need is independent devices that don't call home. But that's not profitable.
I cant speak to whether they are actually sequencing samples, but it'd be a pretty blatant lie if they're not.
So maybe a bunch of human dna gets mixed up with bacterial dna - I dunno really - but it really doesn't matter because of the way they put back together.
Considering they offer the raw data on a 500GB hard drive in a FASTAQ format (https://us.dantelabs.com/pages/faq), I doubt it's just SNPs.
How easy is to get your DNA sample contamination-free by isolating nuclei from dead cells in saliva?
and
"SNP genotyping results for saliva derived DNA (n = 39) illustrated a 98.7% concordance when compared with blood DNA. In conclusion, when compared with blood DNA and tested on the DMET array, saliva-derived DNA provided adequate genotyping quality with a significant lower number of SNP calls. Saliva-derived DNA does perform very well if it contains greater than 31.3% human amplifiable DNA." [2]
[1] https://www.researchgate.net/publication/309849150_Saliva_is...
Sequencing will run to a certain coverage (or depth), aligning multiple fragments so that each nucleotide is sampled e.g. 30 times, which would weed out contaminants. Saliva is generally considered 'good enough' (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3497576/)
They might just use an anti-microbial agent in their collection kit to prevent growth of bacteria until it reaches their lab (plus you don't want bacterial nucleases to fragment your DNA).
Using any genomics is an 'admission' of preexisting condition as a genomic guilt. This $200 is proof you know what diseases you might contract, and this serves as information asymmetry that insurance companies have illegalized.
If I could have this done, say, in Russia or China, and the results destroyed after completed, I'd consider it.
I know that half of my genetic map has already been infiltrated with my mom using 23andme. I felt quite violated when my mom did that. I certainly had no choice. And yet, half of me is sitting in some database, where I have absolutely no say.
And my dad's dead, so no non-degraded DNA there. I guess if my sister ever has the money and want, then more of my map will be known.
The reason I don’t just assume 1x is much worse is George Church is involved in both Nebula and Veritas, and I get the sense he wouldn’t be involved in something if it was worthless.
Correct - the coverage number is an average across the sequenced parts of the genome. In some areas the coverage will be much higher, in others much lower.
Importantly, what is commonly called 'whole genome sequencing' is not really that. There remains ~5-8% of the genome that is (almost) impossible to sequence with current technology, and as such has no coverage. Areas with lots of repeats, centromeres, etc.
But the parent is a bit wrong, it's not worthless. The reason people sequence at <1x is because you can still statistically derive a lot of useful info especially about larger variation - long segments of the genome that are deleted or duplicated etc can be inferred, and there is a technique called "imputation" which means if you measure one part of the genome well you can usually predict the nearby parts with surprising accuracy.
1x is basically 3.3Gb
Maybe if I’m asking the question I’m not ready for the answer? Haha
People also spend money on things like handwriting analysis, personality questionnaires and measuring people's skulls to try to work out what they are like; genomics is at least somewhat accurate and based on reality!
In practice, you'll get told things like "You have endurance muscles, not fast muscles" and "You probably have blue eyes" (mirrors are more accurate!) and "You have twice the chance of getting this specific type of cancer" (twice a very small number is still a very small number).
It's most unfortunate.
In Stranger Visions I collected hairs, chewed up gum, and
cigarette butts from the streets, public bathrooms and
waiting rooms of New York City. I extracted DNA from them
and analyzed it to computationally generate 3d printed
life size full color portraits representing what those
individuals might look like, based on genomic research.
Working with the traces strangers unwittingly left behind,
the project was meant to call attention to the developing
technology of forensic DNA phenotyping, the potential for
a culture of biological surveillance, and the impulse
towards genetic determinism.
The forecast of Stranger Visions came true. Just 2 years
later Parabon NanoLabs launched a service they called DNA
"snapshot" to police around the US. For more examples see
Identitas and read about their collaboration with the
Toronto police.
http://www.identitascorp.com/http://www.theverge.com/2014/7/20/5916661/the-most-advanced-...
https://www.amazon.com/Change-Agent-Daniel-Suarez/dp/1101984...
I thank whoever it was that recommended this book to me some years ago, here on HN. Also, if anyone knows of other books like these, please share!
I recommend campaigning for it, assuming you’re in the USA.
https://en.m.wikipedia.org/wiki/Genetic_Information_Nondiscr...
As I recall from purchasing health insurance in Germany (mandated, and the public option wasn’t available to me as a freelancer), there was price discrimination based on pre-existing conditions, which is currently illegal in the US. The monthly cost was capped at over 600€ Basispreis per individual (this was a couple of years ago), and if you hadn’t visited a hospital for a couple of years and were a healthy man in your 30s and chose to conceal any persitant issues you had been diagnosed with, you could find an offering for closer to 200€. The determination of how pre-existing conditions would affect the price was completely arbitrary and unreasonable, and a significant deterrent to seeking timely care. I’m not aware whether you could choose to not disclose genetic conditions and have your insurance be valid.
(I find it interesting how a lot of folk and religious stories about supernatural sound like memories of civilization with technology comparable to ours.)
I think about this a lot. It has led me to see that level of technology doesn't always linearly increase. At times there appear to have been great regressions proceeding from times of incredible technology. An easy example for my mind are the pyramids in Egypt, or the countless and nauseatinly intricate temples of India. I feel like there have existed technologies that we haven't yet been able to detect or understand.
Biological surveillance is the means by which biological
science is used to track, monitor, analyze, and turn
bodies into data. It is the extraction of DNA and microbes
from our skin, nails, hair and body fluids. It is the
analysis of identifying body parts like faces,
fingerprints and irises. It is the tracking of life itself
by body heat, pulse, perspiration, and involuntary
movement. It is the vulnerability we each face every day
by the very situation of being human, by simply having a
body. Biononymous.me fosters molecular resistance through
the creation of a community to openly discuss, research,
and develop potential solutions through art, science,
technology, policy, and theory.The same genome sequenced two times - how many differences between two sequences? Please, don't tell me it is none.
Sensitivity = true-positive-rate = 0.997.
Precision = 0.997 = #true-positives / (#true-positives + #false-positives) = true-positive-rate / (true-positive-rate + false-positive-rate) = 0.997 => true-positive-rate + false-positive-rate = 1 => false-positive-rate = 0.003. [1]
That seems like a very high error rate, about 10 million errors in the three-gigabase genome, and 100 thousand errors in the 30-megabase exome (protein-coding regions.) That might be an acceptable rate for population-level analysis if the errors are sufficiently uncorrelated, but I wouldn't want to be making decisions on the basis of it for personalized medicine. For comparison, here's a rough estimate that an individual human genome has 2-3 million SNPs [2].
I thought you could do better than that with 30x coverage, so I might be misinterpreting them, somehow. Or maybe they're using an unconventional sequencing technology which is cheaper but less accurate.
[0] https://us.dantelabs.com/products/whole-genome-sequencing-wg...
[1] Equations given here: https://en.wikipedia.org/wiki/Sensitivity_and_specificity
There's no simple answer to your question as it depends on many things - sequencing technology used, library prep and coverage to name a few.
Generally, it's not far from none when aligning short reads to a high-quality reference genome. Provided there's sufficient coverage and a majority of reads covering a particular nucleotide don't have a error at that position, than the correct answer will be given. Errors creep in due to things like systemic errors in library prep (such as a PCR error), and very low coverage over particular loci due to weird AT/GC content, meaning errors are harder to correct for. Repetitive regions can cause issues for short read alignment too, but coding regions generally aren't that repetitive.
$200 is very cheap for WGS - guessing it would be at the low end of the accuracy range, as they can't be sequencing to great depth (presumably).
https://s3.amazonaws.com/dantelabswebsite/Dante+Labs+Genome+...
EDIT: Unless I'm reading the results wrong
Most of the world's whole genome sequencing is done using the Illumina platform. This service is using the BGI platform, which is arguably higher quality than Illumina. Our lab has data showing the error rate with BGI is about 1/6 the error rate of Illumina.
Yes, there are some even better sequencing technologies out there, such as PacBio, which provides longer reads capable of sequencing slightly more of the genome, and the error rates are constantly improving. However, these technologies are much more expensive.
VCF files (the most effective and easy-to-use format) are easily downloadable from your account. FASTQ and BAM files are sent via a 500 GB Hard Disk. To order the Hard Disk, click here