Absolutely embarrassingly parallel. Thing is, I don't know what kinds of problems Xcelerate is interested in working on, I don't know what projects are given preference on Titan, and I don't even know if I'm in the right order of magnitude in my estimate.
There are easier ways to make it more complex. For example, enumerate the first 15 atoms instead of the first 13. I believe this requires coordination between the enumeration nodes in order to reduce duplicate generation, but I don't know that subfield and labeled graph enumeration gets further away from hands-on chemical work. And stage 2 requires a combinatorial approach taking maybe 1,000x more time than the first stage. (That's a jazz hands level of waviness - I really don't know.)
I once did work in molecular dynamics. In fact, I was one of the co-authors of NAMD. That's a rather intently explored field, and progress in it can feel rather incremental. What I proposed is one that no one has explored, to my knowledge. There's maybe 50 people in the world who would be able to do this work now, if they had time and hardware access (which they don't), and perhaps ~5,000 people who would be affected by knowing the result; mostly to know if they could do certain searches on the public web.
So it's one with little competition and where the results would be immediately publishable. Not bad for a semester project, I think! I think it would be a good Master's thesis. And a lot different from the usual set of MD, docking/screening, folding projects that have consumed massive amounts of supercomputer time for the last three decades.
Given that Xcelerate was talking about porting some code to the GPU, I suspect that that person already expect to "spend more time writing the program and the data analysis than the actual runtime." When I was doing MD work, I expected to spend about 6 months in simulation time and about 2 years in analysis for a PhD. The previous generation to me in the group had to build their own hardware before doing simulations, so I though that it's usually the case that the non-simulation time is longer than the simulation time. :)
You said: "I don't really see this problem as a high priority. pharma has bigger problems than to probe remote websites to find out what their competitors might be interested in."
So they should only work on the big problems that everyone else is doing? What about smaller problems which might help solve big problems?
For example, 10 hours ago here on HN, mtgx posted a link titled "Patents Are Making Us Lose The Race Against Antibiotic-Resistant Bacteria" (to http://www.techdirt.com/articles/20130110/09590621628/world-... ), which reports that the "World Economic Forum's 8th Global Risks Report" suggests that "Rather than today's monopolistic hoarding [of data], what we need is more sharing of [pharmaceutical] knowledge."
Let us suppose that they are correct, and we need to have more sharing of some knowledge between different companies, perhaps with "public or philanthropic funding to incentivize academic collaboration."
One of the questions you could ask, if you had more information about what goes on inside of the companies, is: what's the overlap of the chemical space being tested by the different companies? Are they too similar? A recent paper on the topic - and one of the few such papers - is the recent "Big pharma screening collections: more of the same or unique libraries? The AstraZeneca–Bayer Pharma AG case". (See http://pipeline.corante.com/archives/2012/12/06/four_million... for a summary.)
In it, they conclude that there is a "low overlap between both collections in terms of compound identity and similarity."
They did it by using 2D ECFP4 fingerprints. A fingerprint is very much like a Bloom filter, where bits are set based on chemical features. It's one way to get information about identity and similarity. It's not strictly reversible, but it's leaky. For example, fragments can have certain characteristic patterns in the ECFP4 fingerprint, which can be used to infer some of the original structure.
The InChI hashes are another way to test if two data sets contain identical structures. Two pharmas may want to compare them before during a pre-competitive collaboration, in order to determine the overlap between their two collections. What is their risk model? Can they outsource the evaluation knowing that no information can be leaked? Or is there an unexpectedly high leak rate which means that the hashes must never leave the internal network, and should be restricted to only a few trusted people?
Right now, we don't know that answer. We believe that it reveals little information. I'm not so sure about that.
If they don't reveal information, then certain new types of discussions can occur, which might (if the World Economic Forum is correct) help lead to the development of new drugs.
Or, if it does reveal information, then it might help characterize how leaky that information actually is. Eg, it takes 1 month on an Exacycle machine to identify 1% of the structures with fewer than 15 atoms, and 10 years (estimated) to identify 1% of the structures with 17 atoms. That gives a usable risk model that doesn't currently exist.