Man places his genome in public domain, on Github
manu.sporny.org
manu.sporny.org
One major problem with developing a "Google for the human genome" is that we don't actually understand how most of the genes (coding) and noncoding regions in our DNA actually work or interact with each other... except at a very basic level for a very limited set of genes.
There are genome browsers out there already that came out of the human genome project and work in that direction. One example: http://huref.jcvi.org/
It's by no means all of his SNPs, either -- each person has around 3 million actual SNPs (variations from the reference genome), and 23andme just chooses a million sites that could be the location of a SNP to look at, most of which won't actually be points of variation for most people.
So, 23andme is only looking for common SNPs you might have. If you have a rare SNP you're interested in, or if you're a researcher trying to analyze the effects of an uncommon SNP, you're out of luck with 23andme data.
Even though this isn't a genome sequence, there is potential for interesting analyses if lots of people release their 23andme data. (I believe 23andme use the same SNPs for every user).
Additionally, this data could be radically improved if his phenotype was also included. Just because we know that the marker says "AA" without the correlating information of "blond hair" doesn't tell us whether "AA" is important for hair color.
What you can mine with this type of information is the correlation between the markers themselves: if rs1001 = AA, then rs2002 = {GG(85%), TT(10%), GT(5%)}. This is where community software could definitely benefit from more data.
If you are interested in helping create a "Google for DNA", drop us a line at SeqCentral.com
The authors have not only released their genetic information into the public domain (http://www.genomesunzipped.org/data), but also developed a custom genome browser (http://www.genomesunzipped.org/jbrowse), have an API, and a github repo for code they will release (https://github.com/genomesunzipped/genomesunzipped).
These are early days in personal genomics, so it's great to see others jumping in. Hopefully they all do so with some awareness, and folks like Genomes Unzipped do a great job in creating that awareness, and never forgetting that there is difficult, evolving science behind our understanding.
1. http://www.genomicslawreport.com/index.php/2010/03/30/pigs-f...
"Eyelids now close in proper way. Fixes issue #42."
EDIT:
BTW. Any ideas how to get continous integration working with it?
I really like your question. Personally I value stories, creations, etc. in which authors make multiple layers of "jokes" or references. Unfortunately, I'm not good enough to know what letters I'm actually changing :(.
Until that time agent-based modeling is the best we have :)
"These are the 57 public genomes. They are from real people who've chosen to share their data to help all of us learn more about our genomes."