I also maintain our website, so if you see anything that needs fixing, drop me a line :)
I also maintain our website, so if you see anything that needs fixing, drop me a line :)
I'd suggest getting involved in any lab doing any form of computer work. Lab meetings are a great place to find out what problems people are having, and most of the time you'll be able to use your skills to solve them. After taking on a few projects, people will start to realize what you can and can't do, and they'll start to give you a lot of related work.
Good luck!
This can't be overstated. The shear amount of data available for processing is huge. Google-scale. I'm not very familiar with GWAS studies (I'm more familiar with high-throughput sequencing), but these are O(N^m) scale problems (where m>=2).
Where I work, filesystem I/O is the rate-limiting step for most of my (next-gen sequencing) experiments...
However, we just got our own instrument, so this will definitely be an issue, but our University knows a thing or two about dealing with big data (http://kb.iu.edu/data/avvh.html).
I've only dealt with a few GWAS style datasets, and the next-gen stuff dwarfs the GWAS data in terms of size. But when looking for linkages between variations, we're still talking more time than the universe is old level of calculations for more than 3 combinations. Which is really scary, because like you said, all the genetics people are going to be using sequencing for most things from here on out, so its like you have complexity on top of complexity...
OK, now that I've thought about it, the easiest way around this probably is throwing money at hardware (more disks) or optimizing the processing. However, this is only in the case of a true disk I/O bottleneck. If you're optimizing correctly then the disks should be reading 8GB blocks directly into memory and the CPU should be spitting them right out again. At the very minimum you should be using an optimized file system with large pages enabled in your kernel.
"What do you do for a living?"
"Exomes."
-spits drink out-
Regarding massive information, I analyze a lot of the data we generate in order to optimize our delivery platform. Since there are few labs that do this kind of work, all the data I work with comes from us.