I spent a reasonable amount of time pitching this space. I think we need more brilliant people working on hard problems like this.
That being said, I think you might be a bit off the mark and I'm going to share my story and the evidence before me. (Feel free to reach out to me, I'm happy to help anyone trying to tackle problems in this space).
A blog post I wrote several years ago about this kind of thing: www.engineersf.com/2015/07/04/do-we-need-a-human-data-projecthdp/
Our Stab At It
My cofounder and I went around pitching over 100 MDs and disease researchers telling them we had the full medical records of tens of thousands of patients(hint: We didn't).
What They Told Us
What they told us is that the data isn't all that useful for research purposes because it's inaccurate to some extent, but as well the ability to compare and juxtapose patient cases is just really tough. Ron Shigeta @rshigeta on twitter quickly told us the queries they would have to run would be several pages long even for research purposes.
How The Data Could Be Valuable to MDs and Researchers
The REAL pervasive utility in a large number of records is to find the Gene Regulation Pathways, and as humanity we do that through clinical trials. The MDs and Researchers wanted to use it to recruit patients for trials. Fast forward to today and we've focused all our efforts on patient recruitment and screening.
There's another company that's also a YC company that went down a different route and focuses on diagnoses. It might be worth checking them out: humandx.org. (Stands for Human Diagnosis Project)
As well cancer is often highly mutated....ummm...
http://blogs.sciencemag.org/pipeline/archives/2008/09/08/the...
The example you've outlined might be a bad one for the simple reality that cancer is so complicated. Cancer is a cell to cell battle.
Evidence 1: http://blogs.sciencemag.org/pipeline/archives/2011/04/07/mor...
Evidence 2:
http://blogs.sciencemag.org/pipeline/archives/2011/04/05/so_...
The Quote for me that really catches my attention in the pipeline blog:
"Recent work from Bert Vogelstein’s group at Johns Hopkins (with a host of collaborators) and from the CGA itself now show that there are an average of 63 mutations in pancreatic cancer cells, and 47 in glioblastomas, two of the nastiest tumors around. The first impulse might be to think “Great! Plenty of drug targets to go around!”
But hold on. For one thing, even though these mutations are surely not all equal, the fact that there are so many makes you wonder about whether attacking any one of them alone can make much of a difference. And different patients can have varying suites of those mutations, so it’s difficult to imagine that going after just one or two of those targets will be enough to treat a majority of cases. This work follows up on earlier studies in other tumor lines, all of which seem to point in the same direction: patients who are currently classed as having the same type of cancer really don’t.
This won’t come as a surprise to most oncologists, who have seen for themselves the widely varying responses to current therapies. The challenge is to figure out what these various changes mean, and how to classify patients to give them the best therapy. It’s not going to be easy. Just doing the math on the possible interactions of several dozen mutations with a list of possible treatment regimes is enough to make you pause. The hope is that most patients will fall into broad categories, which will line up, more or less, with broad categories of treatment. But it’s not going to be a good fit, most likely, and even getting those approximations to work is taking a lot of time and effort. (Just think back about how long you’ve been hearing about the wonderful new age of personalized medicine. . .)"
I would highly recommend talking to more doctors. Happy to help in any way possible.
Onwards & Upwards.