With a large enough sample pool, we'd be able to correlate features with obscure genes, wouldn't we? Am I missing something fundamental?
It seems like now is the time to get started.
With a large enough sample pool, we'd be able to correlate features with obscure genes, wouldn't we? Am I missing something fundamental?
It seems like now is the time to get started.
As someone who works in the field, the lack of phenotypic data available accompanying genome sequences is currently one of two major roadblocks in advancing genomic science (money is the other problem, of course).
At best, gene sequence data like you describe could provide rough probabilities for some health conditions. That's as likely to lead to over-treatment for a non-problem as catching a real problem early.
Sure, there would be some cases where everyone who has a certain sequence winds up with the same disease, and it'd be good to discover those early in life. (Ex: my wife was born with Wilson's Disease, which causes liver failure by early adulthood 100% of the time without treatment.) But those cases would probably be very few compared to the number of people who get freaked out by things that might become an issue, but probably won't, and the cases where the correlation is too weak to provide a warning, but a health problem occurs.
Siamese cats can have white fur but dark paws and face, it's because of a temperature sensitive mutation enzyme.
Tortoiseshell cats are usually female, and the patterns in their fur are a result of X-chromosome inactivation happening in random clumps of cells.
I think the genes is more like static code. One can clone it easily. RNA is like the stack, RSS, Register states - expression of the genes.
DNA is like a class, RNA is the object that is instantiation of the class.
BTW, the double helix structure is very much like full unit test coverage for the bio-programming.
You may also want to gather more detailed data that the medical history usually conveys, such as sub-clinical conditions and other non-pathological differences.
Nature vs Nurture.
You'd find a huge sampling error because of epigenetics. The same genes in different people don't always kick in and do something.
However, finding out which gene + which life-style == bad stuff might still work out as a possibility.
Even those might not be repeatable in the future - I grew up running around in leaded gasoline land and the same genes will probably never have to deal with so much lead in the air in my kids. Does it matter that I might have special tolerance genes which protect my brain lead fumes?
The study would throw up interesting results, but not "feature -> genes", but maybe "genes -?-> features" (necessary but not sufficient).
I'm skeptical that, just now that we know of epigenetic effects, they will invalidate all previously existing roughly-Mendelian genetics. It's not like people's eyes change from brown to blue if they're stressed in childhood.
So, you wouldn't necessarily find huge sampling error, you'd find sampling error proportionate to the epigenetic effect on that particular gene. Is that large? I suspect not for many genes we'd be interested in.
>Does it matter that I might have special tolerance genes which protect my brain lead fumes?
I don't think this is how epigenetic effects work.
It' so wide spread to be considered a belief on its own: http://www.theallium.com/biology/us-government-to-officially... ;)
What you describe would be useful, but the scope of its usefulness is likely far less than you might think. Basically, you'd only be capturing some long term effects that have very high effect sizes / strong correlations. The sheer complexity, variability, and size of the dataset means that you won't be able to pull out subtle interactions, even with a large number of individuals. Barring fundamental advances in GWAS methods, you're looking at huge expense for very little benefit.
Now, collecting and organizing disparate datasets that are already being generated is worthwhile. Comparison is difficult due to the different methods used, but a good investigator will be able to do a lot of in silico work to, at least, do some preliminary work on hypotheses. On a personal note, this is actually a lot of fun because you can often come across from very exciting hints to follow up on, all with a quick feedback cycle conducive to 'flow'.
Ultimately, and this gets back to science fundamentals, this would amount to a big fishing expedition. Good science will always be strongly hypothesis driven. Good bioinformatics is built upon good datasets, and that only comes from very careful hypothesis consideration, solid methods, and a good understanding of the biology involved.
High throughput phenotypic screens which test the cellular affects of inserting particular variants could help, as they would prune the number of potential variants, but the false negative rate is high.
Genomic determinism may not exist in a useful sense in the current era.
https://en.wikipedia.org/wiki/Dunedin_Multidisciplinary_Heal...
The US military got stung very badly in Desert Storm I with people coming home sick with no way to see what went wrong. They decided not to repeat that mistake.
Disclaimer: I interviewed at 23andme a few years ago, and have purchased their services for my entire family.