Nebula Genomics – First to offer consumer anonymous sequencing
nebula.org
nebula.org
It wasn't long before they figured out who I was and placed me within my family tree. My fake name now lives among near and distant relatives I was not aware had signed up themselves or their parents/grandparents. They know who I am, who my siblings and cousins and aunts and uncles are, etc. This was always going to happen as soon as I sent them my sample.
I never believed my anonymity trick would truly work, I just wanted to make it sufficiently difficult for when 23andme inevitably sold out, got gobbled up, or turned evil. I learned what I wanted from the service, and have only logged in once a year or so since to see if they updated any findings or disease studies.
While I truly appreciate the concept of bringing privacy and anonymity to this field, it's worth considering we are all quite easy to identify using these samples.
You might as well add "hacked" to that list given recent events.
I accepted that and did it anyway, taking steps to at least not be directly associated with my sequence, even if my association can be inferred or derived later. My main concern is that their testing would identify something which in the future would be a "pre-existing condition" and get me denied medical care, but there is certainly a long list of other possible consequences.
At this point I don't trust any company or agency that collects and uses data, or the promises made in any privacy policy, but I also don't lose any sleep over it.
Yes, as long as they have the data. If a company would process the sample, send me a thumb drive of my information, and not retain a copy, that data can't leak because it doesn't exist.
Unfortunately this is just one step away from a blog post where the CEO apologizes for letting down their customers by keeping copies of all data in an unsecured s3 bucket that was downloaded in its entirety by a 13 year old "hacker".
This is why I won't use any genome sequencing service that has a bunch of ancillary services attached (eg. analyzing your ancestry, or figuring out what diseases you're at risk for), and you have to request deletion of data. The fact they provide such services means that your data is getting automatically uploaded to the cloud, probably resulting in multiple copies to different systems/databases/vendors. Even though you can theoretically request deletion, all those copies means there's a non-negligible chance that there's a copy lying around in a decommissioned s3 bucket that they didn't delete. If they service promises sample -> sequencing machine -> lab computer -> [PGP encrypted email/mailed CD], that cuts the risk considerably.
There's nothing logically impossible about such a service, and I'd trust it modulo actual red flags. Too bad afaik nobody's offering it. Once they're archiving their copy I just don't see how they can credibly promise privacy in the longer term.
Last I looked it didn't seem really practical to just buy your own sequencer.
1. Most people don’t care about the privacy aspect
2. People who already got a test from 23andme, Ancestry, etc are unaddressable
Surely, it’s not that costly to delete data? The only reason to keep this data is for ulterior motives like monetizing.
Nebula actually did use to let you download your data and tell them to delete it. When I was looking last year, though, they'd moved to some new model (which I assume this post was about).
I bring this up because Nebula is clearly making a play for anonymity which, in my mind, is strange. Perhaps deceitful. The thing about genetic data, especially WGS, is that it's hopelessly not anonymous. Sure, it can be protected, but given the world we live in (hacks, data brokers, etc.), it's the type of data that will likely get out into the wild. This past week at 23 and Me is a case in point. The comment from another poster here about whether they delete it or not is an important one. I suspect they don't; for a while Nebula was touting a setup where they held onto your data and you could license it for use by pharmaceutical companies or researchers and receive payment. Not sure if that's still a thing. Either way, I find the notion of "anonymity" here very dubious, even if you pay with Monero or similar, use a PO box, etc. But that's kinda the reality – these genomics companies aren't especially open about a several key facts when you get your genome sequence. Another example is that they don't make clear that when you get your genome sequenced you are effectively unmasking the genome of your extended family.
As vatys explains in his longer comment: You are just as anonymous as the least anonymous person related to you.
Then to the privacy policy:
> Nebula will store your Personal Information as long as your Account is open, unless you make a request for us to delete all or any of your Personal Information prior to the closing of your Account as described in this Policy. If you decide to close your Account, then Nebula will automatically destroy all Personal Information related to your account, including User Data, Survey Data, and Genetic Data.
So, my reading of all of this is aligned with how you and vatys have read it, too.
Probably not
I'll probably delete everything off their site in the event it looks like they might get sold but for now I figure it's fairly well guarded and they keep most of it offline anyways.
Also, I can't find anywhere in the Nebula materials describing the chain-of-custody of submitted DNA, & jurisdiction(s) of the sequencing labs.
For example, if provided samples are sent outside the US for sequencing – are they? – I'd have even less confidence they'll be kept secret from local authorities.
You cannot request for this to be deleted or destroyed by law.
I've not seen it reported, & when I last checked (long ago), 23andme seemed to allow a customer to decide whether their sample was retained beyond the current round of tests.
And, if it were a current requirement, wouldn't this Nebula offering be prohibited in California?
https://stanforddaily.com/2018/10/17/me-asl-23-and-me-is-not...
It’s a California “lab quality” law.
Instead, it's said to be kept for at most 10 years, & even that is in a paragraph describing what happens if a customer specifically chooses optional storage!
Further, if you're referencing 23andme's own privacy/terms documents, those also repeatedly pledge that samples are irreversibly destroyed upon request.
So while I can still believe your original claim might be true – California requires many dumb & anti-privacy things, sometimes in dishonest ways! – the links you've forwarded provide more grounds to doubt your claim than support it.
If it's truly a state legal requirement that overrides a user's explicit discard request for a full 10 years, that should be easy to document with a clear and authoritative link, shouldn't it?
Those all seem to be misinformation without support from, and even in contradiction of, public/authoritative references – despite the assertion that these are a matter of "law".
Unfortunately, I was the only one in my family who understood this and the rest were too afraid to get the test done even though people in my family have died of this disease.
A few years ago I tried to get my full genome from nebula but they messed up my order really badly and it took over a year to get my money back. But it looks like they might be cleaning up their act and the $250 price is really tempting.
Since our disorder is polygenic, there’s still a lot of questions I have so I’m probably going to get a nebula run as well.
The interesting this with the data is that it isn't your raw genome, it's the entire data from all the sequencing runs of 30 or whatever base pairs (I got my scan done at the 30x level which is supposed to mean that on average every base pair is read 30 times). So it's not like you can just read your DNA directly in one string (not sure why I want to do that but I like the idea of a single file containing just all my nucleotides in sequence). The next step for me, when I go back to it, is to get on the forums for those above tools and ask how to do the things I want, there are a LOT of subcommands.
https://igv.org/doc/desktop/#UserGuide/reference_genome/
Edit: Since you use Nebula's service, have you emailed them to ask what they recommend for genome viewing and analysis?
I wonder how many of these companies were just his name tacked on for credibility vs him being deeply involved in the product development of direction of the company.
All in comparison to the obvious solution, send in a sample, receive the resulting data. Plus some offline analysis software that you can buy a license for and maybe update from time to time.
They do use statistics in their reports. One of the numbers they give with each report is where you are in the population for that particular report. E.g. 85th percentile for this thing in the population.
Every year the company would print an update to the disease database on more CDs, and make them available for a relatively small price, probably exclusively by mail.
Obviously there's some aspects of this that suck, but it's interesting to see how it's so much better for privacy, because data hoarding wasn't viable at scale or as a business.
* Two, because a genome is > 700MB, so it won't fit on a single disc. How annoying the size limit of a CD would've been if you had to solve this problem at work!
I imagine a couple of floppies a year might've been cheaper in 1995 than a CD.
I would hope it compresses well (lots of repeating stuff), though I have no practical experience.
It's a matter of luck whether your sequencing fails though so I will try again when they offer truly anonymous sequencing (my last attempts were eponymous).
Their website has been saying for at least two years that Anonymous Sequencing is coming. So take their product launch timeline with a grain of salt. In any case true anonymity and the price of 175usd are really impressive if they iron out the logistics.
> Nevertheless, re-identification risk in the wild does not appear to be especially high. While we observe a success rate as high as 25%, this is only achieved when the genomic dataset is extremely small, on the order of 10 individuals. In contrast, success rate for top 1 matching drops quickly and is negligible for populations of more than 100 individuals. Moreover, it should be kept in mind that this result assumes that we can predict the phenotypes perfectly.
"Requires [paid] Nebula Membership".
I see zero reason that I should need to pay them a subscription when all I want is a one-off product. Sure, I can see the use in being alerted when new genetic diseases are discovered, but that should be my choice. It's frustrating that everything is becoming subscription based these days.
Given sufficient time, this information will be useful for the progress of science.
In some sense, this is part of a genetic adaptive mechanism. With a sufficiently tight-band public genetic release with full health history, it is possible that future science will bias to my genes.
Perhaps some of mine will be ones people desire and others not. But they will optimize to the ones that are studied. That means my genes will live on slightly preferentially.
Which, of course, means that I should encourage everyone else to keep theirs private.
Amusing thought. But I think we will all be for the better with our genomes public. It's okay if others won't submit theirs. We will submit ours and may humanity benefit from it.
Genome sequencing has become so cheap in recent years (and will be much cheaper in the next few years) that everyone who wants it will simply get it. Against the background of personalized medicine, this is also very useful. (My background: CEO of a company for personalized cancer therapy through WGS)
https://www.bundesgesundheitsministerium.de/en/en/internatio...
It's better than having your name slapped into it, but let's be honest here unless you have all your material and tools inside your home with no Internet connection, all these things are NOT anonymous by definition.
Case in point: 23andMe
All it takes are a few related (even distantly related) samples and a skilled observer can quickly narrow down where on a genetic tree a sample came from - without knowing anything about the person at all.
[1] https://www.newyorker.com/magazine/2021/11/22/how-your-famil...
1) You aren't sending them naked DNA which can be modified in a test tube, you're sending them epithelial cells that they extract the DNA from.
2) Any changes you make would mostly not be complete, and they're generating your reference genome from a consensus (to account for sequencing errors and DNA damage). This increases the odds that any randomization you do is basically swamped by the consensus algorithm. If the changes are not swamped then your decryption depends on knowing exactly which changes are from your encryption algorithm, and exactly which are from their consensus algorithm.
3) Bisulfite conversion, methyltransferases, general deamination, and the like make known changes to your genome that can easily be fully accounted for during analysis. There's no easily done "random" cryptographic technology that you can do on your DNA. And even if there was, it probably wouldn't be doable without already knowing your specific sequence.
[0]: https://sequencing.com/ [1]: https://sequencing.com/our-difference/privacy-forever
What does that even mean? And do they completely destroy the sampled data from all their systems incl. backups? Can’t find that anywhere.
If not: is it really anonymous?
There’s no plausible reason to use a blockchain instead of a traditional database, and many possible risks/downsides. So either the company is incompetent, or they’re not really using a blockchain and their marketing is misleading, but either way it leaves a bad taste in my mouth.
Also cool to note that this service sequences your entire genome and lets you download the raw data. Which is not something 23 and Me offers.
You can open the data in an open source genome browser and read the raw information that encodes you. Which I just think is the coolest
It took scientists ~18 years to sequence the first genome, I know that was ~15 years ago but they did have huge budgets.